# Get AI chat session information Source: https://docs.prophecy.ai/administration/getting-help/copy-session-info Copy your session info from an AI chat to share with support You can copy your session info from an Agent chat session to share with support or engineering when troubleshooting issues. This information is helpful for: * Reproducing the exact state of your conversation. * Diagnosing issues in your specific environment. ## Steps 1. In your Agent chat, open chat history. 2. Click **Copy Session Details**. 3. The session info will be copied to your clipboard. ## Example session info When pasted, the copied information will look similar to this: ```json theme={null} { "sessionId": "a8c06824-32d5-4ee2-9979-10145f286234", "chatId": "datatransformation_94_1763_dev_default", "hostname": "app.prophecy.io", "commit": "e5740ed9aefff5f8363e82d5a6b235ce52a4f339", "timestamp": 1754659872604, "version": "4.2.8.0" "projectId": "31709", "userEmail": "user.drew@prophecy.io" } ``` # Understand fabric diagnostics Source: https://docs.prophecy.ai/administration/getting-help/diagnostics Troubleshoot fabric issues using diagnostics Troubleshooting Prophecy fabrics is easy with built-in diagnostics. The descriptions are designed to help users to independently identify and resolve issues. When creating or connecting to a fabric, Prophecy automatically tests for connectivity. This feature helps users to determine whether the issue lies within Prophecy itself or in other components of the data ecosystem. ## Diagnostics error codes | Error Code | Symptom | Provider | Cause | Resolution | | ---------- | ------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `10000` | `Is missing from the classpath` | [Databricks](/data-engineering/fabrics/spark-provider/databricks/databricks) | Prophecy Library(Scala) is incorrect. You're probably using thin jar. | Use assembly `jar(${scalaFatJarName})` in the library section of the fabric settings. | | `10001` | `DRIVER_LIBRARY_INSTALLATION_FAILURE` | [Databricks](/data-engineering/fabrics/spark-provider/databricks/databricks) | Prophecy Library(Scala/Python) is incorrect. Databricks could not install it. | Please provide the valid library path in the fabric. | | `10002` | `object prophecy is not a member of package` | [Livy](/data-engineering/fabrics/spark-provider/livy/) | Prophecy Library(Scala) is incorrect. | Please ensure that the library path exists and you're using the assembly `jar(${scalaFatJarName})`. | | `10003` | `cannot be added to user sessions` and `prophecy_libs` | [Livy](/data-engineering/fabrics/spark-provider/livy/) | Prophecy Library(Python) is incorrect. | Please ensure that the library path exists and you're using correct `file(${pythonPLibName})`. | | `10004` | `for method` and `too many arguments` | [Livy](/data-engineering/fabrics/spark-provider/livy/) | Prophecy Library(Scala) is incompatible. | Please use the correct `version(${Globals.prophecyLibsVersion})` in the library section of fabric settings. | | `10005` | `No module named` and `prophecy` | [Livy](/data-engineering/fabrics/spark-provider/livy/) | Prophecy Library(Python) is incorrect. | Please provide the valid library path in the fabric. | | `10006` | `illegal start of simple expression` | [Livy](/data-engineering/fabrics/spark-provider/livy/) | Python version in livy/hadoop is incorrect. | Please make sure you have python3 there. | | `10007` | `IncompatibleClassChangeError` | [Livy](/data-engineering/fabrics/spark-provider/livy/) | Prophecy Library(Scala) is incompatible with your Spark version. | Please use the correct assembly `jar(${scalaFatJarName})` in the library section of the fabric settings. | | `10008` | `"FileNotFoundException` and `prophecy_libs"` | [Livy](/data-engineering/fabrics/spark-provider/livy/) | Prophecy Library(Python) path does not exist. | Please ensure that the file exists as per the path in the library section of the fabric settings. | | `10009` | `503 Service Temporarily Unavailable` and `LivyRestClient` | [Livy](https://livy.apache.org/docs/latest/rest-api.html) | Livy service is down. | Please make sure the livy service is up before executing this command. | | `10010` | `SQLNonTransientConnectionException, rds.amazonaws.com` or `Unable to instantiate, HiveMetaStoreClient` | [Unity Catalog](https://docs.databricks.com/en/resources/supported-regions.html#rds) | Databricks cluster can't access RDS service. | Please ensure that the cluster can access to the same region's RDS endpoint as documented [here](https://docs.databricks.com/en/resources/supported-regions.html#rds). | | `10011` | `UnauthorizedCommandException` and `This execution contained at leas` and `disallowed language` | [Unity Catalog](https://docs.databricks.com/en/resources/supported-regions.html#rds) | Shared cluster in unity catalog does not allow Scala commands. | Please use this cluster with Python Pipeline. | | `10012` | `UnauthorizedCommandException` and `This execution contained at leas` and `disallowed language` | [Databricks](https://docs.databricks.com/en/administration-guide/users-groups/index.html) | This cluster does not allow `${pipeline's language}` command. | Please check with the Databricks workspace administrator to provide the execution access to `${pipeline's language}` language. | | `10013` | `javax.net.ssl.SSLHandshakeException` and `PKIX path building failed` | Livy / [EMR](https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-security.html) | Certificates provided in EMR cluster's security configuration are wrong. | Please ensure that EMR cluster's security configuration is using correct certificates. | | `10014` | `HostUnreachableErrorCode_10014` | [Databricks](https://docs.databricks.com/en/administration-guide/users-groups/index.html) | Unable to reach Databricks endpoint. | Make sure the workspace is active and reachable. | | `10015` | `HiveMetastoreNotEnabledErrorCode_10015` | Hive Metastore | We were unable to write execution metrics because Hive Metastore is not enabled on your Spark. | Please enable Hive Metastore on Spark, or disable execution metrics in Prophecy. | | `10016` | `AuthenticationFAiled_10016` | [Livy](/data-engineering/fabrics/spark-provider/livy/) | Authentication failed. Wrong or no auth credentials were provided. | Make sure correct auth credentials are provided. | | `10ZZZ` | `Library installation attempted on the driver node of cluster and failed` | [Databricks](/data-engineering/fabrics/spark-provider/databricks/databricks) | Databricks runtimes 16.4 and later default to Scala 2.13. | Update Prophecy to version 4.2.0.1+ and update ProphecyScalaLibs for compatibility with Scala 2.13. | # Get in touch Source: https://docs.prophecy.ai/administration/getting-help/get-in-touch We're here to help. Whether you need technical support, want to explore Prophecy's features, or have feedback on our documentation, you'll find the right way to reach us below. ## Self-service Access guides, tutorials, and reference materials. Connect with other users, ask questions, and share knowledge. ## Sales Learn more about Prophecy, pricing, and our Enterprise Edition. Request a personalized product demonstration to see Prophecy in action. ## Support Prophecy provides technical support to customers with an active Prophecy subscription. To contact support, your email address must be registered as an **authorized support contact** for your company. Prophecy may reroute requests submitted by non-authorized contacts to your admin for validation. [Go to the Support Portal β†’](https://prophecy.zendesk.com/) [Email the Support team β†’](mailto:support@prophecy.io) For detailed instructions on how to submit a support request, see [Support requests](/administration/getting-help/submit-request). To learn more about or add authorized support contacts, reach out to our [Sales team](mailto:sales@prophecy.io). Users on any free version of Prophecy are not eligible for support. Please refer to the self-service resources to troubleshoot any issues. ## Documentation feedback To suggest improvements to this documentation, email our [Documentation team](mailto:docs@prophecy.io). # Download HAR files Source: https://docs.prophecy.ai/administration/getting-help/har-file Download HAR files to help Prophecy troubleshoot If you need help from Prophecy Support, it can be helpful to provide HAR files that contain additional information about a certain action you want to debug. This can be especially helpful if the issue involves connections to an execution environment, as the HAR file includes information like timing of network connections and failure messages from API calls. ## Steps To capture the HAR file: 1. Open the page in Prophecy where the problem persists. 2. Right click on the webpage and select **Inspect** (or equivalent). 3. Open the **Network** tab. 4. Refresh the page. You should see websockets created in the Network tab. 5. Perform the action(s) that you want to capture. 6. Export the HAR file. 7. Upload the HAR file to your support ticket on Zendesk. These steps may vary among browsers. Luckily, [Zendesk provides documentation](https://support.zendesk.com/hc/en-us/articles/4408828867098-Generating-a-HAR-file-for-troubleshooting) that outlines these steps for different browsers. # Send logs to Support Source: https://docs.prophecy.ai/administration/getting-help/prophecy-details How to download logs and send for support ## Collect diagnostic data To help our team investigate and resolve your issue faster, include diagnostic data in your support request. Use the table below to find what information to collect based on your issue. | Issue | What to include | | -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Can't attach pipeline to cluster | Connection logs from [Prophecy](./prophecy-details), Spark cluster [configuration](./spark-cluster-details), and [connectivity check](./spark-cluster-details#connectivity-check) results. | | Pipeline fails during execution | Pipeline [logs](./prophecy-details), Spark [configuration](./spark-cluster-details#spark-configurations), and [connectivity check](./spark-cluster-details#connectivity-check) results. | | Spark application errors | Pipeline [logs](./prophecy-details) and Spark [driver logs](https://docs.databricks.com/en/compute/troubleshooting/debugging-spark-ui.html#driver-logs). | ## Send runtime logs If you are having trouble running a pipeline, you can download the [runtime logs](/data-analysis/development/runs/runtime-logs). Our Support team can review these logs to help troubleshoot your issue. Runtime logs ## Send connection logs If you have trouble connecting to your fabric, send connection logs to Prophecy with the corresponding error [code](/administration/getting-help/diagnostics). To retrieve the connection log, open the cluster connection in the top right corner of your project and click the **Copy** button. Connection logs # Security Source: https://docs.prophecy.ai/administration/getting-help/security Learn about Prophecy security practices Prophecy is SOC 2 compliant and safeguards data at every level across the product stack. Whether you use our multi-tenant SaaS platform or a Dedicated SaaS deployment, Prophecy is designed to meet the rigorous security requirements that are standard for enterprises. Read more details on Prophecy's security and compliance posture at our [Security Portal](https://security.prophecy.io/). ## Framework The following is the framework we employ to ensure security with Prophecy. ### Product Prophecy incorporates robust product security features to safeguard user access, data integrity, and compliance. * **Authentication**. Authentication methods, including Single Sign-On (SSO) and Multi-Factor Authentication (MFA), ensure only authorized individuals can access the platform. * **Authorization**. Role-Based Access Control (RBAC) enforces a "least privilege" model, granting users only the permissions they need. * **Auditing**. Detailed auditing capabilities log all user and admin activities for complete transparency and accountability. * **Data protection**. Data is protected through encryption at rest and in motion, with compliant key management practices to safeguard sensitive information. While Prophecy code is stored in Git repositories, your data is not. ### AI and LLMs Prophecy takes a thoughtful approach to AI Large Language Model (LLM) security. Prophecy ensures that AI functionality remains transparent, secure, and trustworthy. * While AI is used to make suggestions, it is never used to directly transform data or make business decisions. * Prophecy **does not** store or send your data to any third-party large language model (LLM) providers. Instead, Prophecy uses rich metadata to construct its [knowledge graph](/data-analysis/ai/knowledge-graph/knowledge-graph). As a result, Prophecy can interface with LLM providers while keeping your data private. * Prophecy Data Transformation Copilot is classified as a "minimal risk" AI application, in accordance with the EU Artificial Intelligence Act. Using an LLM with Prophecy is optional. ### Deployments Here's the full range of deployment options: * **SaaS**. Our highly secure and scalable default offering that employs multi-tenant architecture with multiple layers of logical isolation. Runs on Prophecy's AWS VPC. * **Dedicated SaaS**. A single-tenant deployment option for organizations that prefer more isolated infrastructure deployed on a dedicated Prophecy AWS or Azure VPC. With these flexible options, Prophecy ensures every organization can adopt the deployment model that best suits their security and operational requirements. ### Operations Operational security is a key pillar of Prophecy's approach. The platform is built on hardened infrastructure designed to withstand external threats. Regular penetration testing ensures that vulnerabilities are identified and addressed promptly, while continuous vulnerability scanning and management further strengthen defenses. These practices ensure that Prophecy remains secure, reliable, and resilient against emerging threats. ## Network configuration If you or your organization uses a firewall, VPN, or proxy, Prophecy might not work as expected. Review the following table to help you troubleshoot any issues. | Configuration | Description | | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | WebSocket connection | Prophecy uses a WebSocket connection while talking to certain backend services. Modify your setup to allow this WebSocket connection. | | Whitelist URLs | If you are using SaaS, you must whitelist the following URLs in your network: | | Git providers | Prophecy project code is stored in Git repositories. If you use Git providers within private networks behind firewalls, you must add the Prophecy Control Plane IP address `3.133.35.237` to the private network allow-list or the Git provider [allow-list](https://github.blog/2019-12-12-ip-allow-lists-now-in-public-beta/). | | Databricks connection | If you limit Databricks network access, you must add the **Prophecy Data Plane IP address** `3.133.35.237` to the Databricks allowed [access list](https://docs.databricks.com/security/network/ip-access-list.html). If using Databricks OAuth, you need to ensure network connectivity between your browser and the Databricks workspace. | # Send Spark cluster details Source: https://docs.prophecy.ai/administration/getting-help/spark-cluster-details Helpful Spark cluster configurations to send to Support Available for [Enterprise Edition](/administration/platform/editions) only. There are helpful Spark cluster configurations and a connectivity check that you can send to us via the Prophecy [Support Portal](https://prophecy.zendesk.com/) for troubleshooting. ## Spark configurations Two ways to access the configurations: * Browsing the Spark UI * Running a notebook ### Configurations in the UI You can access your Spark cluster configurations directly from the Spark UI. Please send screenshots of each configuration if possible. | Configuration to Send | Example | | ------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------ | | Overall cluster configuration (e.g., Spark version, Databricks runtime version, UC dedicated or UC standard) |
Cluster configuration example
| | Cluster JSON (edited to remove any private or sensitive information) |
Cluster JSON example
| | Libraries installed on the cluster |
Cluster libraries example
| | Init scripts run on the cluster. Include the script itself if possible. |
Cluster init scripts example
| | Output of attaching cluster in a notebook. You may need to duplicate the tab and try attaching the same cluster in the duplicate tab. |
Notebook attach to cluster example
| ### Run a notebook For those who prefer to use code, create a notebook (example below) and send the output via the Prophecy [Support Portal](https://prophecy.zendesk.com/). Replace the workspace URL, personal access token, clusterID, and API token as appropriate. ```python theme={null} # Databricks notebook source import requests #Get Databricks runtime of cluster # Get the notebook context using dbutils context = dbutils.notebook.entry_point.getDbutils().notebook().getContext() # Retrieve the Databricks runtime version from the context tags runtime_version = context.tags().get("sparkVersion").get() # Print the runtime version print(f"Databricks Runtime Version: {runtime_version}") # Get Spark version spark_version = spark.version print(f"Spark Version: {spark_version}") #Get the installed libraries and access mode details of the cluster # Replace with your Databricks workspace URL and token workspace_url = "replace_with_workspace_url" token = "replace_with_token" cluster_id = "replace_with_cluster_id" # API endpoint to get info of installed libraries url = f"{workspace_url}/api/2.0/libraries/cluster-status" # Make the API request response = requests.get(url, headers={"Authorization": f"Bearer {token}"}, params={"cluster_id": cluster_id}) library_info=response.json() print("Libraries:") for i in library_info['library_statuses']: print(i) # API endpoint to get access mode details url = f"{workspace_url}/api/2.1/clusters/get" # Make the API request response = requests.get(url, headers={"Authorization": f"Bearer {token}"}, params={"cluster_id": cluster_id}) cluster_access_info=response.json() print(f"Cluster Access Mode: {cluster_access_info['data_security_mode']}") ``` ## Connectivity Check Open a notebook on the Spark cluster and run the following command. Replace the Prophecy endpoint. ```python theme={null} import subprocess command = 'curl -X GET "https://customer_prophecy_url/execution"' output = subprocess.check_output(['/bin/bash', '-c', command], text=True) print(output) ``` ```scala theme={null} %scala import sys.process._ val command = """curl -X GET "https://customer_prophecy_url/execution"""" Seq("/bin/bash", "-c", command).!! ``` This command tests the reverse websocket protocol required by Prophecy to execute pipelines on Spark clusters. Please send the output from this command in the Support Portal. **We look forward to hearing from you!** # Prophecy status page Source: https://docs.prophecy.ai/administration/getting-help/status-page Monitor the health of your Prophecy environment and subscribe to incident notifications The Prophecy Status Page provides real-time visibility into the health and availability of your Prophecy environment. Use the status page to: * Check if your Prophecy environment is currently operational. * Track ongoing incidents and maintenance windows. * View historical uptime and past incident details. * Subscribe to notifications so you are alerted immediately when something changes. Prophecy Status Page ## Status indicators Each component is assigned a status indicator: | Status | Description | | -------------------- | ------------------------------------------------------------------------------------------------------------------------- | | 🟒 Operational | All systems are running normally. No action needed. | | πŸ”΄ Major Outage | A significant portion of the platform is unavailable. Incident is open and being worked on with high priority. | | 🟣 Under Maintenance | Scheduled maintenance is in progress. This is a planned activity communicated in advance. Temporary disruption may occur. | | βšͺ Unknown | Status cannot be determined. This may indicate a monitoring gap. Contact support if this persists. | ## View incidents When an issue affects the platform, Prophecy posts updates on the status page. Incidents typically progress through the following states: 1. **Investigating** - The team has been alerted and is actively looking into the issue. Not all root causes are known at this point. The status page will be updated as information becomes available. 2. **Identified** - The root cause has been found. The team is now working on a resolution. An estimated time to fix may be shared at this stage. 3. **Monitoring** - A fix has been applied and the team is watching to confirm the issue is fully resolved. Services may still show as degraded during this period. 4. **Resolved** - The incident is closed. All affected services are back to normal operation. A post-incident summary may be shared with affected customers. Incident updates are displayed in chronological order and include information about impact, mitigation efforts, and resolution status. ## Subscribe to notifications You can subscribe to be notified by email when incidents are created, updated, or resolved. To subscribe: 1. Open the status page. 2. Scroll to **Subscribe to status updates** (1). 3. Enter your email address. 4. Click **Subscribe**. 5. Confirm the subscription using the link sent to your inbox. After confirmation, you will receive notifications for future status changes and incidents. Subscribe to status updates Confirmation links expire after 24 hours. If your link expires, start a new subscription. If you have confirmed but are not receiving notifications, do the following: * Check your spam or junk folder. * Add [updates@statuspage.betterstackupdate.com](mailto:updates@statuspage.betterstackupdate.com) to your contacts or safe sender list. * Verify that you subscribed to the correct status page URL. ## Manage subscriptions To stop receiving notifications, click **Unsubscribe** in any status page email. To change the email address on your subscription, unsubscribe from the current address and create a new subscription with your new email address. (You cannot edit email addresses on existing subscriptions.) ## Scheduled maintenance The Prophecy team may schedule planned maintenance windows for upgrades, infrastructure changes, or other operational activities. These are communicated on the status page in advance. * Scheduled maintenance events appear on the status page before the window begins. * Subscribers receive email notifications ahead of the maintenance start time. * During maintenance, affected services display the Under Maintenance status (🟣). * Once maintenance is complete, all services return to green (Operational) and you will receive a confirmation notification. If you have time-sensitive workflows that could be affected by a maintenance window, please reach out to your Prophecy support contact as soon as you see a maintenance announcement so we can plan accordingly. ## Contact support The status page reports platform-wide service health and incident information. For account-specific issues, integration questions, or anything not reflected on the status page, contact Prophecy support: * Submit a support ticket through the [Zendesk Support Portal](/administration/getting-help/submit-request). * Email [support@prophecy.io](mailto:support@prophecy.io) For P1/production-blocking issues, contact your dedicated Prophecy SRE or Customer Success Manager directly. # Support requests Source: https://docs.prophecy.ai/administration/getting-help/submit-request Understand the support request process This page explains how Prophecy customers with support included in their subscription can open and manage support cases. If your organization does not have support access, you can use self-service resources such as documentation and tutorials. ## Prerequisites To open support tickets and new cases, you need to be registered as an authorized support contact for your company. For more information, reach out to our [Sales team](mailto:sales@prophecy.io). Users on free versions of Prophecy are not eligible for support. Please refer to self-service resources to troubleshoot any issues. ## Submit a support ticket * Open the [Support Portal](https://prophecy.zendesk.com). * Click **Submit a Ticket**. Zendesk support portal homepage Complete the support request form with the information below. Ticket submission form **Subject**: Enter a short summary of the issue. **Description**: Provide detailed information about your issue. Include: * A description of the issue, detailing expected vs. actual behavior. * Any steps taken so far to troubleshoot the issue. * The workspace URL where the issue occurred. * The pipeline, dataset, or other resource involved. **Urgency**: Select the option that best reflects the impact of the issue. **CC (optional)**: Add team members to the ticket. Users must exist in Zendesk to be added as CC. **Current Prophecy Version (optional)**: Include your Prophecy version to help diagnose version-specific issues. **Attachments**: Upload logs, screenshots, or other relevant files (up to 10 MB). For information on what to include, see the [Send logs to Support](/administration/getting-help/prophecy-details) page. Allowed file types: PDF, PNG, JPEG, GIF, TXT, LOG, HAR, JAR. Click **Submit** to create your support ticket. ## View your support tickets * Open the [Support Portal](https://prophecy.zendesk.com). * Click **View all tickets**. You can filter the requests by status, including: * Any (shows all requests) * Open * Awaiting your reply * Solved In addition to your own requests, you can also view: * Requests that you are CC'd on * Organization requests (any ticket submitted by other users in your organization) Zendesk support portal requests
view ## Case severity Prophecy uses severity levels to prioritize issues based on business impact. | Level | Description | Initial Response | Coverage | | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------- | ------------------------- | | **Severity 1**
Urgent | Production use is blocked/broken without a workaround. | \< 1 hour | 24/7 every day | | **Severity 2**
High | Major functionality is impaired. System is functional but in a degraded or restricted state. | \< 4 hours | 24/7 every day | | **Severity 3**
Normal | Non-critical functionality is impacted or a development-only issue occurs. | \< 1 business day | 9 AM - 6 PM business days | | **Severity 4**
Low | General questions or feature requests with no production impact. | \< 1 business day | 9 AM - 6 PM business days | If a ticket requires additional attention, you may escalate by: * Updating the ticket with new details. * Contacting your customer success manager. # Troubleshooting Source: https://docs.prophecy.ai/administration/getting-help/troubleshooting View common issues and solutions encountered in Prophecy Prophecy projects can run into occasional issues related to version control, deployments, connections, or runtime environments. This page collects common problems and their solutions. If you are working on a project and don't see changes that you previously made, several issues could be the cause: * You are on a different branch than the one containing your changes. * Another user may have overwritten your changes or reverted commits. * Multiple tabs were open, and one version overwrote another. This happens in the Normal Git Storage Model. The Simple Git Storage Model's single-user mode prevents this. Learn more in [Version conflicts](/data-analysis/development/versioning/conflicts). If you cannot make edits in the pipeline canvas: * The pipeline may have been imported from another project, which makes it non-editable. Pipelines are only editable from the original project. * The pipeline might be [read-only](/data-analysis/development/versioning/conflicts) to prevent merge conflicts when others are editing the same project. If this is the case, you'll be able to see who is editing the pipeline at that time and can request control of the pipeline. A pipeline run may fail with generic or unknown errors. * Check runtime logs for detailed error messages. * Confirm that all inputs, expressions, and connections are correctly configured. A project may not appear in your workspace if you are not part of the assigned team. * Verify your team membership for the project. * Ask an admin to update team access if necessary. Even if a pipeline runs successfully, the target table may remain empty. * Confirm that the target connection is valid and not expired. * Check that the write mode settings are correctly configured. Output may not match expectations due to incorrect pipeline logic or write mode settings. * Verify write mode settings (Append vs. Overwrite). * Check filters and expressions in the pipeline logic. Scheduled pipelines may fail to trigger if the project containing the schedule is unpublished. * Ensure the project is published and the schedule is active. Custom SQL may fail if syntax is incorrect or the dialect is unsupported. * Verify SQL syntax and ensure the dialect used matches your SQL warehouse dialect. Pipeline errors β€” especially with Source and Target gems β€” may occur if credentials or tokens have expired. * Update connection credentials and any saved secrets. Previously successful pipelines may fail if connections expire, schemas change, or other edits occur. * Update expired connection credentials. * Check Source/Target gems for schema changes. * Review pipeline version history to identify changes made by others. Prophecy cannot retrieve datasets if the authenticated identity does not have access on the connection side. * Confirm that your credentials grant access to the required datasets in the origin data source. If a Table gem with an explicit target schema resolves to `._.` instead of `..
`, the project is missing the auto-generated `generate_schema_name` macro. With the macro removed, dbt's built-in version concatenates the fabric default schema with the table target schema. * Open the project sidebar and check the **Functions** section for an entry named `generate_schema_name`. * If it is missing, restore it. Learn more in [Schema name resolution in SQL projects](/data-analysis/development/extensibility/generate-schema-name). # Audit logs Source: https://docs.prophecy.ai/administration/management/audit-logs How Prophecy generates, stores, and shares audit logs This page describes how Prophecy generates, stores, and shares audit logs for Prophecy deployments. Setting up audit logs requires collaboration with Prophecy. [Contact Prophecy](https://www.prophecy.io/request-a-demo) to: * Export audit logs on demand. * Configure automatic syncing to your own storage. * Set a custom retention period for stored logs. ## Storage location Prophecy stores audit logs in the same cloud platform as your deployment: * For AWS deployments, audit logs are stored in Amazon S3. * For Azure deployments, audit logs are stored in Azure Blob Storage. * For Google Cloud Platform deployments, audit logs are stored in Google Cloud Storage. ## Audit event reference When audit logs are enabled for your Prophecy deployment, they capture the following information: * User interactions with the Prophecy UI. * GraphQL API calls made to the control plane. * Actions performed directly on the execution plane, such as fabric, connection, deployment, pipeline run, and schedule operations. The following tables list the audit events that Prophecy logs, organized by entity type. Prophecy uses GraphQL for control plane API operations. Request and response parameters may vary depending on where you call the query. ### Fabric | Query | Description | Request Parameters | | ------------------- | ---------------------------- | ---------------------- | | `fabricDetailQuery` | Get Fabric Details | `["id"]` | | `addFabric` | Add a Fabric | `["name", "ownerUid"]` | | `updateOwnedBy` | Update Team owing the Fabric | `["id","targetUid"]` | | `userFabricQuery` | Get all Fabrics for User | `["uid"]` | ### Project | Query | Description | Request Parameters | | -------------------------- | --------------------------------------------------- | ----------------------------------------------------------------------------- | | `addProject` | Add a project | `["name","forkMode","language", "ownerUid", "mainBranchModificationAllowed"]` | | `getDetails` | Get Details of a Project | `["projectId"]` | | `project` | List all projects for User | `["uid"]` | | `teamProjectAvailable` | Available Projects for that Team | `["uid", "language"]` | | `addProjectDependency` | Add a dependency Project to Current | `["projectId", "DependencyProjectUid"]` | | `updateProjectDependency` | Update dependency Project to a new released version | `["projectId", "DependencyProjectUid", "ReleaseTag"]` | | `removeProjectDependency` | Removed an added dependency | `["projectId", "DependencyProjectUid"]` | | `projectDependenciesQuery` | List all project Dependencies | `["projectId"]` | | `projectReleaseStatus` | Gives Status of last Release for given project | `["projectID", "statuses"]` | | `projectSyncFromGit` | Status of Git sync of project | `["uid"]` | | `releaseProject` | Release a Project | `["branch", "message","version","projectID", "CommitHash"]` | | `gitFooter` | Details for Git for commit/branchNAme etc | `["projectID"]` | | `addSubscriberToProject` | Add Subscriber to a Project | `["uid", "teamId"]` | | `projectBranches` | List of available branches for this project | `["projectId"]` | | `cloneProject` | Created clone of current project | `["uid", "name", "teamUid", "copyMainBranchReleaseTags"]` | ### Pipeline | Query | Description | Request Parameters | | ---------------------- | ------------------------------------ | -------------------------------------------------------------------------------- | | `addPipeline` | Add a new Pipeline | `["name", "branch", "ownerId", "doCheckout"]` | | `tableQueryPipeline` | Lists all pipelines for project | `["projectId", "sortOrder", "sortColumn"]` | | `tableQueryPipeline` | Lists all pipelines for User | `["uid", "sortOrder", "sortColumn"]` | | `pipelineDetailsQuery` | Get Details of Pipeline | `["Uid"]` | | `clonePipeline` | Cloned a Pipeline | `["branch", "sourcePipelineId", "targetPipelineName", "ownerUid", "doCheckout"]` | | `addSubgraph` | When Subgraph is added to a Pipeline | `["mode", "name", "language", "ownerUID"]` | | `addUDFBulk` | UDFs added to a Project | `["udfs.name","udfs.description", "projectUID"]` | | `removeUDFBulk` | UDFs removed form a project | `["uids"]` | | `getSubgraph` | Get Subgraph by given Id | `["uid"]` | ### Job | Query | Description | Request Parameters | | ------------------------------------ | ---------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | | `addJob` | Add a Job | `["name", "branch","fabricUID", "scheduler", "doCheckout", "projectUID"]` | | `updateJobConfiguration` | Job configurations are updated | `["emails", "jobUID", "enabled", "onStart", "fabricId", "onFailure", "onSuccess", "clusterMode", "scheduleCron"]` | | `latestJobReleaseByJobIdAndFabricID` | Get Jobs Release by Fabric Id | `["jobUID", "fabricUID"]` | | `jobReleaseByProjectRelease` | Gets Jobs Released by Project ID | `["projectReleaseUID"]` | | `jobQuery` | Get a Job by given Id | `["uid"]` | | `addJobRelease` | Adds a Job released mapping to project Release | `["jobUID", "fabricUID", "scheduler", "schedulerJobUID", "projectReleaseUID"]` | | `tableQueryJob` | list query for Jobs | `["uid", "sortOrder", "sortColumn"]` | ### Dataset | Query | Description | Request Parameters | | --------------------- | --------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | | `queryDataset` | When Datasets are queried from any page | `["uid", "optionalProjectUID"]` | | `addDataset` | Added a new Dataset | `["mode", "name", "ownerUID", "fabricUID", "datasetType"]` | | `addMultipleDatasets` | Add Multiple Datasets | `["names", "ownerUID", "tableNameList", "schemaNameList", "descriptionsList", "schemaAspectList", "databaseNamesList"]` | ### Team | Query | Description | Request Parameters | | -------------- | ------------------------------------- | ---------------------------------------------- | | `addTeam` | Added a new Team | `["name", "adminUid"]` | | `getUserTeam` | Get Teams for a User | `["uid"]` | | `addteamAdmin` | Add a user as Admin | `["teamUid", "userUid", "invitationAccepted"]` | | `user` | List All teams for Users with Members | `["uid"]` | ### User | Query | Description | Request Parameters | | ------------------------- | ------------------------ | ------------------------------------ | | `getUser` | Get User | `["email"]` | | `tableQueryUser` | List query for the User | `["uid", "sortOrder", "sortColumn"]` | | `userAllFabricInfoAspect` | Get User Details | `["uid"]` | | `setPassword` | user Sets a new Password | `["uid", "newPassword"]` | ### Git | Query | Description | Request Parameters | | | ----------------------------- | ------------------------------------------------- | --------------------------------------------------------------------------------------- | - | | `deleteBranch` | Deleted a Branch | `["projectId", "branchName"]` | | | `checkout` | Checkout a new branch | `["projectId" , "branchName"]` | | | `prCreationRedirectUrl` | Pr Creation button clicked | `["to", "from", "projectId"]` | | | `createPR` | Pr Creation button clicked | `["to", "from", "toFork", "fromFork":, "projectId"]` | | | `cleanCommit` | Committed any changes | `["message", "projectId"]` | | | `commit` | Commit button clicked | `["branch", "message", "projectId"]` | | | `pullOrigin` | pull origin branch | `["branch", "projectId"]` | | | `checkGitConnection` | Test Git connection | `["externalUriArg", "pushAccessCheck", "userGitCredsUID"]` | | | `linkSavedCredsToExternalGit` | Linked Saved Creds to a project | `["projectUID", "userGitCredsUID"]` | | | `unlinkExternalGit` | Unlink the saved creds | `["projectUID"]` | | | `branchDivergence` | When user compares two branches for commit screen | `["projectId", "branchName", "baseBranchName"]` | | | `branchInfo` | Gives details of a particular working branch | `["projectId", "branchName", "remoteType"]` | | | `setPrCreationTemplate` | When user Sets PR creation template | `["projectId", "customPrTemplate", "prCreationEnabled"]` | | | `getPrCreationTemplate` | Gets PR creation template | `["projectId"]` | | | `deleteUserGitCreds` | When user deleted saved Git creds | `["uid"]` | | | `linkExternalGit` | Link saved Git creds | `["projectUID", "externalRepoUri", "userGitCredsUID"]` | | | `mergeMaster` | Merge to master branch | `["prNumber", "projectId", "entityConflicts", "projectConflicts", "resolvedConflicts"]` | | ### Transpiler | Query | Description | Request Parameters | | --------------------- | -------------------------------------- | ----------------------------------------------------- | | `transpilerImport` | Transpiler Import started | `["uid"]` | | `addTranspilerImport` | Importing files to Prophecy Transpiler | `["name", "status", "storagePath", "transpilerType"]` | ### Generic | Query | Description | Request Parameters | | -------------- | -------------------------- | -------------------------------------------------------------- | | `removeEntity` | When any entity is removed | `["uid", "entityKind"]` | | `updateEntity` | When any entity is updated | `["uid", "entityKind", "entityFieldName", "entityFieldValue"]` | ## Execution plane events Prophecy audit logs also capture actions performed directly on the execution plane (the orchestrator service running in your environment), not just control plane GraphQL calls. These events are recorded locally on the execution plane first, then forwarded to the control plane in batches (every 30 minutes by default). As a result, there's a short delay between an execution action occurring and its event appearing in the central audit log. Execution plane events are identified by a fixed event type and category rather than a GraphQL query name. ### Authentication and identity | Event type | Category | Description | | ------------------ | ---------- | ---------------------------------------------------------------------------------------- | | `login_federated` | `auth` | A user signed in through a federated identity provider (OAuth, OIDC, or SAML) | | `login_ldap` | `auth` | A user signed in through LDAP | | `login_builtin` | `auth` | A user signed in with a Prophecy-native username and password | | `login_token` | `auth` | A user signed in with a bearer token (reserved for future use) | | `logout` | `auth` | A user ended their own session | | `user_impersonate` | `identity` | A privileged user (for example, a Support role) requested a token to act as another user | ### Fabric | Event type | Category | Description | | ---------------------- | ----------- | ------------------------------------ | | `fabric_create` | `data_flow` | A fabric was created | | `fabric_config_update` | `data_flow` | A fabric's configuration was updated | | `fabric_delete` | `data_flow` | A fabric was deleted | ### Connections and secrets | Event type | Category | Description | | ---------------------- | ----------- | ------------------------------ | | `connection_create` | `data_flow` | A connection was created | | `connection_update` | `data_flow` | A connection was updated | | `connection_delete` | `data_flow` | A connection was deleted | | `fabric_secret_create` | `secrets` | A secret was added to a fabric | | `fabric_secret_update` | `secrets` | A fabric secret was updated | | `fabric_secret_delete` | `secrets` | A fabric secret was deleted | ### Deployment and pipeline runs | Event type | Category | Description | | -------------------------- | --------- | ------------------------------------------------------------------- | | `project_deploy` | `compute` | A project was deployed to a fabric | | `pipeline_run_interactive` | `compute` | A pipeline was run interactively, for example from the Prophecy IDE | | `pipeline_run_scheduled` | `compute` | A pipeline run was triggered by a schedule | | `pipeline_run_sync_api` | `compute` | A pipeline was run synchronously through the REST API | ### Schedules | Event type | Category | Description | | ----------------- | --------- | ---------------------- | | `schedule_create` | `compute` | A schedule was created | | `schedule_update` | `compute` | A schedule was updated | | `schedule_delete` | `compute` | A schedule was deleted | ## Sync data to S3 If your Prophecy deployment is hosted on AWS, you can sync your Prophecy audit logs to your own Amazon S3 bucket. Follow these steps to configure your S3 bucket and grant Prophecy the required access. ### 1. Create the S3 bucket 1. Open the Amazon S3 console and choose **Create bucket**. 2. Enter a **Bucket name**, following the format `prophecy-customer-audit-events-foo`. Replace `foo` with an identifier for your organization. 3. Choose a **Region**. Prophecy syncs to whichever region your bucket is created in β€” there's no region you need to avoid or request special handling for. 4. Complete the remaining setup options as needed, then create the bucket. 5. Set **Object Ownership** to **ACLs disabled (recommended)**. You can apply this setting during bucket creation or by editing bucket permissions after creation. 6. If your bucket uses a customer-managed KMS key, grant the Prophecy role `kms:Encrypt`, `kms:GenerateDataKey*`, and `kms:DescribeKey` in your key policy. Without this, the sync will create successfully but fail when Prophecy tries to write objects. ### 2. Configure bucket permissions for Prophecy 1. In the Amazon S3 console, open your bucket and choose the **Permissions** tab. 2. Under Bucket policy, select **Edit**. 3. Paste the following policy JSON, replacing the placeholders as described below. ```json theme={null} { "Version": "2008-10-17", "Statement": [ { "Sid": "DataSyncCreateS3LocationAndTaskAccess", "Effect": "Allow", "Principal": { "AWS": "arn:aws:iam::##############:role/AWSDataSyncS3BucketAccessCustomerAuditEventsRole" }, "Action": [ "s3:GetBucketLocation", "s3:ListBucket", "s3:ListBucketMultipartUploads", "s3:AbortMultipartUpload", "s3:GetObject", "s3:ListMultipartUploadParts", "s3:PutObject", "s3:GetObjectTagging", "s3:PutObjectTagging", "s3:DeleteObject" ], "Resource": [ "arn:aws:s3:::prophecy-customer-audit-events-foo", "arn:aws:s3:::prophecy-customer-audit-events-foo/*" ] }, { "Sid": "DataSyncCreateS3Location", "Effect": "Allow", "Principal": { "AWS": "arn:aws:iam::##############:user/s3access" }, "Action": "s3:ListBucket", "Resource": "arn:aws:s3:::prophecy-customer-audit-events-foo" } ] } ``` To use this example JSON: * Replace all instances of `prophecy-customer-audit-events-foo` with your bucket ARN. * The two statements grant access for different purposes: the first lets Prophecy's sync role read your bucket's location and read/write objects at transfer time. The second β€” the `DataSyncCreateS3Location` statement β€” grants Prophecy's `s3access` IAM user the `s3:ListBucket` permission it needs to register your bucket as a sync destination in the first place. Both are required. * After applying the policy, contact Prophecy and provide: * Your bucket ARN. * The AWS region. Prophecy will complete the configuration and enable syncing for your environment. # Authentication options Source: https://docs.prophecy.ai/administration/management/authentication/authentication Use your identity provider to sign in to Prophecy Custom authentication is available for the [Enterprise and Express Editions](/administration/platform/editions) only. When logging in to Prophecy, you can either credentials managed directly by Prophecy, or set up SSO. Prophecy integrates with multiple identity providers to let you log in using your external credentials. You can configure SSO under **Settings > SSO**. Only [Prophecy cluster admins](/administration/management/users/access/role-based-access) have permission to view and edit SSO settings. ## Prophecy-managed authentication By default, Prophecy uses **Prophecy Managed** authentication. This option requires no external identity provider. * User accounts are created and managed inside Prophecy. * Passwords are stored securely within Prophecy. * Use this mode if you don't have an external SSO requirement. If you set up SSO after creating users in Prophecy, sign-ins will map to existing users if the sign-in email matches the user email in Prophecy. ## Fabric OAuth To set up OAuth for fabric connection authentication, see [OAuth app registrations](/administration/management/cluster-admin-settings/oauth-setup). # Microsoft Entra ID SSO Source: https://docs.prophecy.ai/administration/management/authentication/azure-ad Sign-in to Prophecy using your Microsoft Entra ID credentials Available for [Express and Enterprise Editions](/administration/platform/editions) only. Prophecy supports **direct OAuth integration** with Microsoft Entra ID (formerly Azure Active Directory). ## 1. Register a new app First, you need to log in to the [Azure portal](https://portal.azure.com/) as an administrator and register a new app. 1. In the Azure portal, open the **App registrations** page. 2. Click **New Registration**. 3. Name it `ProphecyEntraIDApp`. 4. Choose the supported account type: **Accounts in this organizational directory only (`xxxxx only - Single tenant`)** 5. For the Redirect URI, choose **Web** in the dropdown and use: `https://your-prophecy-ide-url.domain/api/oauth/azureadCallback` 6. Click **Register**. ## 2 (Optional): Enable automatic team creation To automatically create new teams in Prophecy via [group mappings](/administration/management/authentication/group-team-mapping), follow these steps. 1. In your Prophecy deployment, set the `ENABLE_AUTO_TEAM_CREATION` flag to `true`. 2. Open the Azure portal. 3. Open the app that you registered in [1. Register a new app](#1-register-a-new-app). 4. Under **Manage**, select **Token configuration**. 5. Select **Add groups claim**. 6. Select the **Groups assigned to the application** checkbox. * To change the groups assigned to the application, select the corresponding application from the **Enterprise applications** list. Select **Users and groups** and then **Add user/group**. Select the group(s) you want to add to the application from **Users and groups**. 7. Click **Save**. These steps are also listed in the [Configure groups optional claims](https://learn.microsoft.com/en-us/entra/identity-platform/optional-claims?tabs=appui#configure-groups-optional-claims) section of the Microsoft documentation. ## 3. API Permission Next, go to **API permissions** on the left-hand side and add this set of API permissions: ![Screenshot 2022-06-13 at 9 57 16 PM](https://user-images.githubusercontent.com/59466885/173400731-acb084df-31a7-4858-b6ba-f395e888e60e.png) ## 4. Certificates and Secrets Then, go to **Certificates and Secrets**, add a new secret, and note down the value of this secret. ## 5. Client ID Finally, click on **Overview** on the left-hand side and note down the Application (client) ID. ## 6. Configure Prophecy to connect with Microsoft Entra ID 1. Log in to Prophecy as an admin user. 2. Navigate to the **SSO** tab of the Prophecy **Settings** page. 3. Under **Authentication Provider**, select Azure Active Directory. 4. Enter the **Client ID** and the **Client Secret** at minimum. 5. Click **Save**. Once you have logged out, you will be able to see a **Login with Azure Active Directory** option. Now, your Azure AD users will be able to login to Prophecy with this option. # Google SSO Source: https://docs.prophecy.ai/administration/management/authentication/google-sso Sign-in to Prophecy using your Google account credentials This setup is applicable to Express and Enterprise Editions only. Google SSO is automatically configured for Free and Professional Editions. Prophecy supports direct OAuth integration with **Google Identity**. ## Configuration steps First, create a new OAuth client in Google. 1. Create an **OAuth client** in Google Cloud Console. 2. Collect the **Client ID** and **Client Secret**. 3. Add Prophecy's **Redirect URI** to the OAuth client. ``` https:///api/oauth/googleCallback ``` Next, add these values in Prophecy. 1. Log in to Prophecy as a cluster admin. 2. Navigate to **Settings > SSO**. 3. Under **Authentication Provider**, select **Google**. 4. Fill in the **Client ID** and **Client Secret**. 5. Click **Save** at the bottom of the page to save your changes. This allows users to sign in to Prophecy with their Google Workspace credentials. # Group-to-team mapping Source: https://docs.prophecy.ai/administration/management/authentication/group-team-mapping Automatic team creation and user role assignment based on identity provider groups Available for [Express and Enterprise Editions](/administration/platform/editions) only. Prophecy supports automatic team creation and user role assignment based on identity provider groups using either SCIM or LDAP. To learn about roles in Prophecy, visit [Role-based access](/administration/management/users/access/role-based-access). ## Overview When a user signs in through an identity provider, Prophecy checks their group memberships and uses naming conventions to: * Create teams in Prophecy * Assign users to those teams * Set the user's roles This applies to both **SCIM** and **LDAP**-based identity systems. ## Standard naming conventions By default, Prophecy supports the following group naming patterns. | Group Name Pattern | Role in Prophecy | | ------------------ | ------------------------------- | | `-user` | Member of the `` team | | `-admin` | Admin of the `` team | | `prophecy-admin` | Prophecy cluster admin | ## Custom naming conventions If your organization uses more complex naming schemes (for instance, with prefixes or suffixes), Prophecy can still infer team names and assign roles appropriately. The following custom patterns are supported. | Group Name Pattern | Example | Role in Prophecy | | ---------------------------------------------------- | ------------------------------------- | ------------------------------ | | `(-)prophecy-cluster-admin(-)` | `corp-prophecy-cluster-admin` | Prophecy cluster admin | | `--admin` | `corp-finance-admin` | Admin of the **finance** team | | `--user` | `corp-finance-user` | Member of the **finance** team | | `--prophecy-team-admin(-)` | `corp-sales-prophecy-team-admin-emea` | Admin of the **sales** team | | `--prophecy-team(-)` | `corp-sales-prophecy-team-emea` | Member of the **sales** team | Prophecy will automatically strip the `` and `` from group names to determine the team name. The prefix and suffix will **not** appear in the Prophecy UI. ## Required configurations To enable automatic team creation using custom naming schemes, you must configure the following environment variables for your deployment. Please reach out to us to set up these configurations. ### For LDAP To use automatic team creation, enable the following flag. ``` ENABLE_AUTO_TEAM_CREATION: "true" ``` This tells Prophecy to dynamically create teams when matching LDAP group names. ### For both SCIM and LDAP To identify and remove prefixes from team names in Prophecy, add the prefix to the `PROPHECY_IDP_TEAMNAME_STRIP_REGEX` variable. ``` PROPHECY_IDP_TEAMNAME_STRIP_REGEX: "prefix-" ``` Replace `prefix-` with your actual prefix (for example, `corp-`, `info-`, etc.). Prophecy will strip this from group names before creating or matching teams. You can use regular expressions for the prefix value. # LDAP authentication Source: https://docs.prophecy.ai/administration/management/authentication/ldap Connect to your organization LDAP directory for authentication Available for [Express and Enterprise Editions](/administration/platform/editions) only. Prophecy can connect to your organization's LDAP directory for authentication. * You can integrate Microsoft Entra ID over LDAP by using the same configuration fields. * Group mapping can be enabled using [group-to-team mapping](/administration/management/authentication/group-team-mapping). To automatically create new teams in Prophecy via [group mappings](/administration/management/authentication/group-team-mapping), set the `ENABLE_AUTO_TEAM_CREATION` flag to `true` in your Prophecy deployment. ## Configuration Here are the basics steps to connect Prophecy with your LDAP directory: 1. Log in to Prophecy as an admin user. 2. Navigate to the **SSO** tab of the Prophecy **Settings** page. 3. Under **Authentication Provider**, select LDAP. 4. Fill out the rest of the information and click **Save**. More information about the available fields can be found below. ### Host and Certs | Parameter | Description | | ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Host | Host and optional port of the LDAP server in the form `host:port`. If the port is not supplied, it will be guessed based on "Disable SSL" and "Start TLS" flags. | | Disable SSL | Required if the LDAP host is not using TLS (port 389). This option inherently leaks passwords to anyone on the same network. | | Skip Certificate Verification | If a custom certificate isn't provided, this option can be used to turn on TLS certificate checks. | | Certificates | Upload trusted Root certs, client certs, and client keys. | ### Binds | Parameter | Description | | ----------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | | Bind Distinguished Name | The distinguished name for an application service account. The connector uses these credentials to search for users and groups. | | Bind Password | The distinguished password for an application service account. The connector uses these credentials to search for users and groups. | | Username Prompt | The attribute to display in the provided password prompt. | ### User Search | Parameter | Description | | ----------------------- | ------------------------------------------------------ | | Base Distinguished Name | BaseDN to start the search from. | | Filter | Optional filter to apply when searching the directory. | | User Name | Username attribute used for comparing user entries. | | ID Attribute | String representation of the user. | | Email Attribute | Attribute to map to Email. | | Name Attribute | Maps to display name of users. | ### Group Search | Parameter | Description | | ----------------------- | ------------------------------------------------------ | | Base Distinguished Name | BaseDN to start the search from. | | Filter | Optional filter to apply when searching the directory. | | Name Attribute | Maps to display name of users. | ### Configured LDAP Groups API You can use the Configured LDAP Groups API to retrieve all config data for your LDAP groups. Example: ``` curl 'https:///api/idp/getAllIDPsConfig' \ -H 'Content-Type: application/json;charset=utf-8' \ -H 'cookie: prophecy-token=' ``` Response: ``` { "data": { "config": [ { "id": "cp_ldap", "type": "ldap", "name": "", "idp": "others", "resourceVersion": "", "idpConfig": { "host": "host-name-here:host-port", "insecureNoSSL": true, "insecureSkipVerify": true, "startTLS": false, "rootCA": "", "clientCert": "", "clientKey": "", "rootCAData": "", "clientCertData": "", "clientKeyData": "", "bindDN": "*****", "bindPW": "*****", "usernamePrompt": "cn", "userSearch": { "baseDN": "dc=example,dc=org", "filter": "(objectClass=person)", "username": "cn", "scope": "", "idAttr": "DN", "emailAttr": "mail", "nameAttr": "cn", "preferredUsernameAttr": "", "emailSuffix": "" }, "groupSearch": { "baseDN": "ou=users,dc=example,dc=org|ou=newusers,dc=example,dc=org", "filter": "(objectClass=groupOfNames)", "scope": "", "userAttr": "", "groupAttr": "", "userMatchers": null, "nameAttr": "cn" } } }, { "id": "cp_saml", "type": "saml", "name": "", "idp": "okta", "resourceVersion": "", "idpConfig": { "caData": "-----BEGIN CERTIFICATE-----\nCERT-HERE\r\n-----END CERTIFICATE-----\n", "emailAttr": "email", "entityIssuer": "issuer", "groupsDelim": ", ", "nameIDPolicyFormat": "persistent", "redirectURI": "https://env-domain/api/oauth/samlCallback", "ssoIssuer": "http://www.okta.com/TOKEN", "ssoURL": "https://SSO-URL", "usernameAttr": "name" } } ] }, "success": true } ``` ## User Matchers This list contains field pairs that are used to match a user to a group. It adds a requirement to the filter that an attribute in the group must match the user's attribute value. # SAML authentication (SCIM optional) Source: https://docs.prophecy.ai/administration/management/authentication/saml Leverage SAML for authentication for Prophecy users Available for [Express and Enterprise Editions](/administration/platform/editions) only. Security Assertion Markup Language (SAML) lets Prophecy delegate user authentication to your identity provider (IdP). System for Cross-domain Identity Management (SCIM) optionally automates provisioning and deprovisioning of users and teams from your IdP into Prophecy. This page describes how to set up SAML and SCIM, which requires configuration in both Prophecy and your preferred IdP. ## Prerequisites Review the following prerequisites. * To access SSO settings, you must be a [cluster admin](/administration/management/users/access/role-based-access) for your deployment. * SAML is available for [Express and Enterprise Editions](/administration/platform/editions). * SCIM is only available for the [Enterprise Edition](/administration/platform/editions). To enable SCIM in your environment, update the `config` in your Prophecy deployment. ## Supported identity providers Prophecy supports the following identity providers (IdP): * Google * Okta * Azure Active Directory (Microsoft Entra ID) * Others (custom) ## Prophecy-specific steps ### Set up SAML To set up SAML authentication in Prophecy: 1. Log in to Prophecy as a cluster admin user. 2. Navigate to the **SSO** tab of the Prophecy **Settings** page. 3. Under **Authentication Provider**, select SAML. 4. Under IdP, select the appropriate identity provider. 5. Fill out the remaining parameters: | Parameter | Description | | ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | | Generate SCIM Token | If SCIM is enabled in your environment, click to create/rotate the SCIM bearer token used by your IdP/SCIM client to provision users and groups. | | SSO URL | The Identity Provider Single Sign-On endpoint. Paste the SSO URL (Location) from your IdP metadata. Prophecy redirects users here to authenticate. | | Certificate | The IdP SAML signing certificate. Prophecy uses this to verify SAML response signatures. | | Skip certificate verification | Disable TLS certificate/host verification when calling the IdP SSO endpoint. | | Entity issuer | The entity issuer you configured in the IdP. Usually, this is the label for your SAML configuration. | | SSO issuer | The SSO issuer. For Azure AD this is the Azure AD Identifier; for Okta, it is the issuer link in the XML. | 6. Click **Save** at the bottom of the page to save your changes. SSO settings for SAML and SCIM configurations ### Set up SCIM SCIM allows Prophecy to automatically provision and deprovision users and teams based on your IdP configuration. To set up SCIM in Prophecy: 1. Ensure that SCIM is enabled in your Prophecy environment. 2. Complete the [SAML configuration](#set-up-saml) described above. 3. Click **Generate SCIM Token** and copy the value of the token. You will use this in your IdP settings later. After completing these steps, you'll create users and groups in your IdP directly. See [Group-to-team mapping](/administration/management/authentication/group-team-mapping) to learn about group naming conventions. ## IdP-specific steps The following sections describe the fields you must configure in your IdP settings to complete the SAML/SCIM setup. ### Azure Active Directory Configure SAML for Azure Active Directory (Microsoft Entra ID) and enable SCIM provisioning 1. Log into AzureAD as an administrator and create a new Enterprise Application like `ProphecyAzureADApp`. 2. In the home page search bar, search for **Enterprise Applications**. 3. Click **New Application > Create your own application**. 4. Give name for the application like `ProphecyAzureADApp`. 5. Choose the radio button **Integrate any other application you don't find in the gallery (Non-gallery)**. 6. Click **Create**. 7. In Manage section on the left, click **Single sign-on**. 8. Choose **SAML** as the Single sign-on method. Now the form for **Set up Single Sign-On with SAML** will open. You'll have to fill out different sections of the form. #### Basic SAML Configuration 1. Provide an Identifier (Entity ID) which is a unique ID to identify this application to Microsoft Entra ID. This will be added to the Entity issuer field in Prophecy. 2. In the same section, configure **Reply URL** and **Sign on URL** as: `https://your-prophecy-ide-url.domain/api/oauth/samlCallback` 3. Click **Save**. #### Attributes & Claims 1. Click **Edit** button and then **Add new claim**. 2. Give **Name** as `email` and **Source Attribute** as `user.userprincipalname`, and click **Save**. 3. Add one more claim by clicking on **Add new claim**. 4. Give **Name** as `name` and **Source Attribute** as `user.givenname`, and click **Save**. #### SAML certificates In the **SAML certificates** section, download `Certificate (Base64)` file to be used while configuring SSO in Prophecy UI. #### Set up ProphecyAzureADApp In the **Set up ProphecyAzureADApp** section, copy `Login URL` and `Azure AD Identifier` to be used while configuring SSO in Prophecy UI. AzureAD config example #### Configure provisioning To set up SCIM in Azure: 1. Open the Enterprise Application you just configured. 2. From the left menu of the application, click **Manage > Provisioning**. 3. From the left menu of the provisioning page, click **Manage > Connectivity**. * Under **Select authentication method**, select **Bearer authentication**. * For the **Tenant URL** field, use the value `https://.prophecy.io/proscim` * For **Secret token**, use the Prophecy-generated token that you copied in the [Set up SCIM](#set-up-scim) steps. * Click **Test connection**. Once the test succeeds, you can save the connection. 4. From the left menu of the same provisioning page, navigate to **Manage > Provisioning**. * Set the **Provisioning Mode** to **Automatic**. * Ensure that the **Provisioning Status** toggle is set to **On**. 5. From the left menu of the same provisioning page, navigate to **Manage > Users and groups**. * Create groups using Prophecy's [naming conventions](/administration/management/authentication/group-team-mapping). * Assign users to those groups. These users and teams should automatically appear in Prophecy. ### Okta Configure SAML for Okta and enable SCIM provisioning 1. Log in to Okta as an administrator. 2. On the homepage, navigate to **Applications** > **Applications**. 3. Click **Create App Integration**. 4. Select **SAML 2.0** and click **Next**. 5. Enter **App Name** as *Prophecy SAML App* and click **Next**. 6. For **Single Sign-On URL**, specify `https://your-prophecy-ide-url.domain/api/oauth/samlCallback`. 7. Select **Use this** for both **Recipient URL** and **Destination URL**. 8. In **Audience URI (SP Entity ID)**, provide a name to serve as the entity issuer ID (for example, `prophecyokta`). 9. Set **Name ID format** to **EmailAddress** from the dropdown. 10. For **Application Username**, select **Email**. 11. Under **Attribute Statements**, add two attributes **name** and **email**. Okta config example 12. Click **Next**. 13. Choose **I'm an Okta customer adding an internal app**. 14. Click **Finish**. The *Prophecy SAML App* is now displayed. #### Download SAML Signing Certificate 1. Navigate to the **Sign On** tab of *Prophecy SAML App* in Okta. 2. Locate the **SAML Signing Certificates** section. 3. Click the download button, as shown in the example below, to download the certificate: Download Okta Cert #### SSO URL 1. In the same **Sign On** tab under **SAML Signing Certificates**, click **View IdP metadata**. 2. This action opens an XML file in a new browser tab. 3. Copy the red-highlighted text in the **Location** section of the XML file and use it as the **SSO URL** in Prophecy IDE. IdP Metadata #### Entity and SSO Issuer 1. Go to the **General** tab, then navigate to the **SAML Settings** section and click **Edit**. 2. Click **Next** to reach the **Configure SAML** section. 3. Scroll to the bottom and click the **Preview the SAML assertion** button. 4. This opens a new browser tab. 5. Copy the highlighted information from the preview and use it as the **Entity Issuer** and **SSO Issuer** in Prophecy IDE. SAML Assertion #### Enable SCIM provisioning To set up SCIM in Okta: 1. Open in Prophecy application in Okta. 2. In the **General** settings, select the **Enable SCIM provisioning** checkbox. Once SCIM is enabled, two new tabs should appear in the application settings: **Provisioning** and **Push Groups**. 1. Navigate to the **Provisioning** tab. 2. In the **To App** subtab, enable the following: * **Create Users** * **Update User Attributes** * **Deactivate Users** This allows Okta to perform the enabled actions in Prophecy. 3. In the **Integration** subtab: * Set the **SCIM version** to `2.0`. * For the **SCIM connector base URL**, use the value `https://.prophecy.io/proscim` * For the **Unique identifier field for users**, type `userName`. This is an Okta-specific value. * For **Supported provisioning actions**, ensure that **Push New Users**, **Push Profile Updates**, and **Push Groups** checkboxes are enabled. * Set the **Authentication Mode** to **HTTP Header**. * Under **HTTP Header**, set the **Bearer** token to the Prophecy-generated token that you copied in the [Set up SCIM](#set-up-scim) steps. * Click **Test Connector Configuration**. * When successful, click **Save**. Now that you have set up the connection, you need to create and activate your group assignments. 1. Open the **Assignments** tab. 2. Create users and assign them to groups using Prophecy's [naming conventions](/administration/management/authentication/group-team-mapping). 3. Navigate to the **Push Groups** tab. 4. Follow the steps in [Enable Group Push](https://help.okta.com/en-us/content/topics/users-groups-profiles/usgp-enable-group-push.htm) in the Okta documentation to activate group provisioning in Prophecy. Once you have pushed groups, the corresponding users and teams should appear in Prophecy. When groups are "active" in Okta, any changes to these groups should automatically sync with Prophecy. # Settings list Source: https://docs.prophecy.ai/administration/management/cluster-admin-settings/cluster-admin-settings Settings visible to cluster admins only Available on the [Enterprise Edition](/administration/platform/editions) only. [Prophecy cluster admins](/administration/management/users/access/role-based-access) have access to additional settings in the **Settings** UI of their Prophecy environment. To access these settings: 1. Click **... > Settings**. 2. Open the **Admin** tab. ## Admin tabs The Admin settings contains various tabs that serve different purposes. | Tab | Description | | ---------- | ------------------------------------------------------------------------------------------------------------- | | Users | View or remove users that exist in your deployment, download Mixpanel data, or upgrade your Prophecy version. | | Teams | View or remove teams in your deployment. | | Fabrics | View fabrics and their teams and authors. | | Config | Update configurations related to alerts, backups, logs, object store, etc. | | Backup | View previous backup information and backup status. | | Monitoring | Review the resource usage of various services in the Prophecy deployment. | | Security | Global authentication settings like Databricks OAuth U2M setup, Kerberos Authentication, etc. | | Logs | Download system logs from a specific time range. Helpful for support cases. | Cluster admin settings page # Configure Anthropic endpoint for Prophecy Source: https://docs.prophecy.ai/administration/management/cluster-admin-settings/configure-endpoint Set up LLM provider for Anthropic Claude in Prophecy. Prophecy allows you to use Anthropic Claude models through different providers by configuring the Transform Agent in settings. Only one provider can be active at a time. Applicable to the [Express and Enterprise Editions](/administration/platform/editions) only. ## Navigate to settings 1. Log in to Prophecy as an administrator. 2. Go to **Settings > Admin > Copilot Settings**. 3. Locate **AI Model Provider Credentials**. 4. Edit the **Transform Agent JSON**. ## Configure Anthropic (public API) Use this option to connect directly to Anthropic using an API key from the Anthropic Console. Apply the following configuration: ```json theme={null} { "transform_agent": { "anthropic_key": "" } } ``` This configuration sends requests directly to Anthropic's public API using your provided API key. ## Configure Azure AI Foundry Use this option to route Claude requests through your Azure AI Foundry deployment. Apply the following configuration: ``` { "transform_agent": { "foundry_config": { "use_foundry": "true", "foundry_api_key": "YOUR_API_KEY", "foundry_base_url": "YOUR_BASE_URL", "foundry_model": "claude-opus-4-5" } } } ``` ### Configuration parameters | Parameter | Description | | ---------------------- | ---------------------------------- | | **foundry\_api\_key** | Your Azure AI Foundry API key | | **foundry\_base\_url** | Your Azure AI Foundry endpoint URL | | **foundry\_model** | The deployed Claude model name | ## Important * Remove `/v1/messages` from the base URL provided by Azure AI Foundry. * The model name must match the deployment name in Azure AI Foundry. ## How it works Prophecy uses the Transform Agent configuration to determine which provider handles Claude requests. Defining either an `anthropic_key` or a `foundry_config` block selects the provider. Only one provider can be configured at a time. # Copilot settings Source: https://docs.prophecy.ai/administration/management/cluster-admin-settings/copilot-settings Configure model provider credentials and model specifications for Copilot Applicable to the [Express and Enterprise Editions](/administration/platform/editions) only. Prophecy's AI capabilities, including Copilot and Agents, are powered by external LLMs. * The [SaaS](/administration/platform/prophecy-deployment#saas) deployment uses a Prophecy-managed OpenAI subscription with GPT-4.1 and GPT-4.1 mini. * [Dedicated SaaS](/administration/platform/prophecy-deployment#dedicated-saas) deployments connect to customer-managed endpoints. In customer-managed deployments, you configure providers and models in **Copilot Settings**. Prophecy uses two AI configuration patterns: * [Agents](/data-analysis/ai/agent/agent) use only the `transform_agent` configuration in **AI Model Provider Creds** to determine how Claude is accessed by Agents. * [Copilot](/data-analysis/gems/copilot/gem-expressions) uses provider credentials and model selection to power features like expression generation. These configurations serve different features, are configured independently, and live in the same settings. ## Navigate to Copilot settings 1. Log in to Prophecy as an Administrator. 2. Go to **Settings > Admin > Copilot Settings**. Copilot Settings includes three subtabs: * [AI Model Providers Creds](#ai-model-providers-creds): Add credentials for LLM providers. * [Available AI Models](#available-ai-models): Define models for Copilot. * [Available AI Speech Models](#available-ai-speech-models): Define models for speech features. All settings must be provided in YAML format and override values in your Kubernetes deployment. ## How Copilot and Agents use these settings Agents use `transform_agent` in [AI Model Providers Creds](#ai-model-providers-creds) Copilot uses: * [AI Model Providers Creds](#ai-model-providers-creds). * [Available AI Models](#available-ai-models). * [Available AI Speech Models](#available-ai-speech-models). These configurations are independent and can be used together. ## Configure AI Model Providers Creds (Agent and Copilot) In the **AI Model Providers Creds** subtab, you provide the credentials to connect to your LLM provider. You can configure both Copilot providers and an Agent provider in the same credentials block. (It is common to use both.) The following YAML example shows the required fields for different LLM providers. ```yaml theme={null} { "transform_agent": { "anthropic_key": "********"}, 'azure_openai': { 'api_key': '********', 'api_endpoint': '********' }, 'openai': { 'api_key': '********' }, 'gemini': { 'api_key': '********' }, 'vertex_ai': {}, } ``` The `transform_agent` block is used only by Agents to configure access to Claude. Copilot does not use this configuration.For details on configuring endpoints for the Agents, see [Configure Claude providers](/administration/management/cluster-admin-settings/configure-endpoint). Other providers in this section are used only by Copilot when selecting models. You can add multiple credentials here, but Copilot will only connect to the models defined in the [Available AI Models](#available-ai-models) subtab. Agents use the `transform_agent` configuration instead. When editing credentials in Prophecy, you will not be able to see previously-entered values. Likewise, once you save your credentials, values will be masked by asterisks `****` as shown in the example above. #### Set up Vertex AI Vertex AI does not require an API key. Instead, you provide a [Google Cloud service account](https://docs.cloud.google.com/iam/docs/service-account-overview) file. Reach out to the [Support team](mailto:support@prophecy.io) with the following information to add Vertex AI credentials to your deployment: * The service account file that can authenticate your connection to Vertex AI. * (Optional) The Vertex AI [region](https://docs.cloud.google.com/docs/geography-and-regions), if it differs from the service account region. * (Optional) The Vertex AI [project](https://docs.cloud.google.com/resource-manager/docs/creating-managing-projects), if it differs from the service account project. ### Available AI Models for Copilot This configuration applies only to Copilot features. Agents do not use these models. In the **Available AI Models** subtab, you define two models: * **Smart model**: Prophecy uses this model for complex tasks. Recommended Model: `gpt-4.1` * **Fast model**: Prophecy uses this model for easy and quick tasks. Recommended Model: `gpt-4.1-mini` You can select models from any of the providers you configured in the [AI Model Providers Creds](#ai-model-providers-creds) subtab. The following YAML example shows how to format your smart and fast model configuration. ```yaml theme={null} { 'smart_model': { 'provider': 'openai', 'model_name': 'gpt-4.1' }, 'fast_model': { 'provider': 'openai', 'model_name': 'gpt-4.1-mini' }, } ``` ### Available AI Speech Models This configuration applies only to Copilot speech features. In the **Available AI Speech Models** subtab, you define different models for speech-to-text and text-to-speech operations. You can select models from any of the providers you configured in the [AI Model Providers Creds](#ai-model-providers-creds) subtab. The following YAML example shows how to format your speech-to-text (`stt`) and text-to-speech (`tts`) model configuration. ```yaml theme={null} { 'stt': { 'provider': 'openai', 'model_name': 'whisper-1' }, 'tts': { 'provider': 'openai', 'model_name': 'tts-1' }, } ``` ### Prerequisites for configuring Copilot providers These prerequisites apply only to Copilot features. To add your LLM provider and model details to Copilot settings: * You must be logged in as a Prophecy cluster admin. * Copilot must already be enabled in your Prophecy deployment. Reach out to the [Support team](mailto:support@prophecy.io) to enable Copilot. # Feature management Source: https://docs.prophecy.ai/administration/management/cluster-admin-settings/feature-management Enable or disable features in your Prophecy deployment Available for the [Enterprise Edition](/administration/platform/editions) only. Feature flags control which capabilities are available to users in your Prophecy deployment. Use them to restrict access to specific platform features, manage fabric types, or toggle AI functionality without having to directly modify your Kubernetes configuration. ## Prerequisites To access feature management, you must be logged in as a Prophecy cluster admin. ## Feature list Access feature management from **Settings > Admin > Feature Management**. Any changes you make to feature flags take effect only after you restart your Prophecy deployment. The following features are available in the Feature Management tab. | Feature | Description | | ----------------------------------------------------- | ------------------------------------------------------------------------------------------------- | | Artificial Intelligence | Enable or disable Copilot in your deployment. | | AI Agents | Enable or disable Agents in your deployment. | | Full Text Search | Enable or disable full text search for your deployment. | | Prophecy Automate | Enable or disable [Prophecy Automate](/administration/platform/architecture) for your deployment. | | Enable Databricks SQL warehouse classic fabrics | When enabled, users can create SQL fabrics with Databricks as the provider. | | Enable Google BigQuery fabrics | When enabled, users can create Prophecy fabrics with Google BigQuery as the provider. | | Enable Snowflake fabrics | When enabled, users can create Prophecy fabrics with Snowflake as the provider. | | Limit fabric and connection creation to cluster admin | When enabled, only cluster admins can create fabrics and data ingress/egress connections. | | SCIM (System for Cross-domain Identity Management) | Enable or disable SCIM for your deployment. | | Transpiler | Enable or disable the transpiler for your deployment. | | Enable Alteryx to Databricks SQL Fabric | Enable or disable the ability to transpile Alteryx workflows into Databricks SQL projects. | While Prophecy provides the feature flags listed above in the Admin settings UI, there are additional features that are set directly in your Kubernetes deployment. Reach out to the [Support team](mailto:support@prophecy.io) if you need any other customization. # Knowledge graph configuration Source: https://docs.prophecy.ai/administration/management/cluster-admin-settings/knowledge-graph-config Configure the V4 agent to use the knowledge graph when AI data access is disabled Available on the [Enterprise Edition](/administration/platform/editions) only. When `AI_DATA_ACCESS_CLUSTER_ENABLED` is set to `false`, the V4 agent can still serve metadata-based capabilities (dataset discovery, schema lookups, and KG search) as long as you have configured the knowledge graph service URL and model. Without this configuration, the agent will fail to call KG tools even when the knowledge graph has been indexed. ## Prerequisites * Kubernetes access to the namespace where Prophecy is deployed * An existing and successfully indexed [knowledge graph](/data-analysis/ai/knowledge-graph/knowledge-graph) ## Configure the knowledge graph service Update the `sql-sandbox-config-map` to specify the Claude model and knowledge graph service URL. If the ConfigMap already exists, append the keys under `data`. Replace `` with your actual namespace. ```yaml theme={null} apiVersion: v1 kind: ConfigMap metadata: name: sql-sandbox-config-map namespace: data: CLAUDE_MODEL: claude-opus-4-5 KNOWLEDGE_GRAPH_BASE_URL: http://knowledge-graph-service..svc.cluster.local:50055 ``` Apply the configuration: ```bash theme={null} kubectl apply -f .yaml ``` Then restart the orchestrator to pick up the changes: ```bash theme={null} kubectl -n rollout restart deploy orchestrator- ``` After restarting, verify that the agent can resolve dataset and schema queries before closing your session. See [Agent behavior when data access is disabled](/data-analysis/ai/knowledge-graph/knowledge-graph#agent-behavior-when-data-access-is-disabled) for expected behavior. # OAuth app registrations Source: https://docs.prophecy.ai/administration/management/cluster-admin-settings/oauth-setup Create app registrations in Prophecy for OAuth setup Available for the [Express and Enterprise Editions](/administration/platform/editions) only. Configure OAuth for Prophecy fabric connections by creating app registrations for supported identity providers. ## Prerequisites Before you configure OAuth in Prophecy, ensure you have: * Cluster admin access to Prophecy. * An OAuth application created with your identity provider. See [Create provider-side OAuth applications](#create-provider-side-oauth-applications). * The client ID from your provider's OAuth application. ## Supported providers Prophecy supports OAuth authentication with the following providers: * **Databricks**: Authenticate with Databricks workspaces. * **Google**: Authenticate with Google Cloud services. * **ID Anywhere**: Authenticate with custom identity providers. * **ID Anywhere (Snowflake)**: Authenticate with Snowflake using an ADFS/ID Anywhere identity provider instead of Snowflake's native OAuth. ## App registration selection If you create multiple app registrations for a certain provider, the selection behavior varies based on the fabric type. | Fabric type | Description | Example | | ------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | | [Prophecy fabrics](/data-analysis/environment/fabrics/prophecy-fabrics) | You can select which app registration to use from a dropdown in the fabric connection settings. If multiple app registrations exist, you can toggle between them. | If you create a Prophecy fabric and configure a Databricks connection, and you have multiple Databricks app registrations, you can select which one to use. | | [Spark fabrics](/data-engineering/fabrics/spark-provider/databricks/databricks) | The default app registration for the provider is **always** used automatically. You cannot change which app registration is used at the fabric level. | If you create a Spark fabric and select Databricks as the provider, the fabric will always use the default Databricks app registration. | ## Create an app registration To add a new OAuth app registration: 1. Sign in to Prophecy as a cluster admin. 2. In the navigation menu, go to **Settings** > **Admin**. 3. Select the **Security** tab. 4. Click **Add App Registration**. 5. Configure the registration settings: | Field | Description | Required | | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------ | | Provider | Select your identity provider (Databricks, Google, ID Anywhere, or ID Anywhere (Snowflake)). | Yes | | Default for Provider | When enabled, this registration becomes the default OAuth configuration for fabrics for the selected provider. The default is always used for Spark fabrics and cannot be changed at the fabric level. | No | | Name | A descriptive name to identify this registration. Useful when managing multiple registrations for the same provider. | Yes | | App Client ID | The client ID from your OAuth application. | Yes | | App Client Secret | The client secret from your OAuth application. | No | | Token Lifetime | Override the default token lifetime set by your provider. | No | | Authorization Endpoint | The authorization URL for your identity provider. | ID Anywhere, ID Anywhere (Snowflake) | | Scopes | Space-separated list of OAuth scopes. Required for ID Anywhere and ID Anywhere (Snowflake). Optional for Databricks and Google to override [default scopes](#default-and-custom-scopes). | Depends | 6. Click **Save**. ### Default and custom scopes Each provider requires specific OAuth scopes: | Provider | Default scopes | Custom scopes documentation | | ----------------------- | ---------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | | Databricks | `all-apis`, `offline_access`, `profile`, `email`, `openid` | [Custom app integration scopes](https://docs.databricks.com/api/account/customappintegration/create) | | Google | `https://www.googleapis.com/auth/bigquery` | [OAuth 2.0 Scopes for Google APIs](https://developers.google.com/identity/protocols/oauth2/scopes) | | ID Anywhere | None (must specify manually) | Consult your provider's documentation | | ID Anywhere (Snowflake) | None (must specify manually) | The relying-party ID your ADFS issues for this audience. It must also appear in Snowflake's `external_oauth_audience_list`. | ## Create provider-side OAuth applications Before adding an app registration in Prophecy, you need to create the corresponding OAuth application with your provider. ### Databricks First, a Databricks [account admin](https://docs.databricks.com/en/admin/index.html#what-are-account-admins) needs to complete the following steps **once** for your Prophecy deployment: 1. On Databricks, navigate to **Account Settings > App connections** in your account console. 2. [Create a new App connection](https://docs.databricks.com/en/integrations/enable-disable-oauth.html#enable-custom-oauth-applications-using-the-databricks-ui) for Prophecy. Ensure that: * Access scopes are set to **ALL APIs**. * The redirect URL contains the following URLs: ``` https:///api/databricks/oauthredirect https:///metadata/oauthCallback ``` 3. This process generates Databricks OAuth Application fields on the Prophecy side. 4. Under Client ID, copy your **OAuth Client ID** for the application, and share it with your Prophecy Cluster Admin. 5. Under Client secret, select **Generate a client secret**. Share it with your Prophecy Cluster Admin. 6. Click **Save**. ### Google Cloud Create an OAuth 2.0 client in Google Cloud Console: 1. Sign in to [Google Cloud Console](https://console.cloud.google.com). 2. Select your project. 3. Go to **APIs & Services** > **Credentials**. 4. Click **Create Credentials** > **OAuth client ID**. 5. Configure the OAuth consent screen if prompted. 6. Select the application type. 7. Add the following redirect URI: ``` https:///metadata/oauthCallback ``` 8. Save the generated client ID and secret for use in Prophecy. For detailed instructions, see [Setting up OAuth 2.0](https://support.google.com/cloud/answer/6158849) in the Google Cloud documentation. ### ID Anywhere For custom identity providers, consult your provider's documentation to: * Create an OAuth 2.0 application or client. * Configure the authorization endpoint. * Define the required scopes. * Generate client credentials. * Set up redirect URIs to point to your Prophecy instance. ### ID Anywhere (Snowflake) Use this provider when you want a Snowflake fabric to authenticate through an ADFS/ID Anywhere identity provider instead of Snowflake's native OAuth. It can point at the same ADFS instance used for a Databricks ID Anywhere registration, or a different one. Each registration's Authorization Endpoint is used as-is and isn't shared between registrations. Before adding the app registration in Prophecy: * Register the relying party (or equivalent) with your ADFS admin, and note its ID; you'll enter this as the registration's **Scopes** value. * Add that same relying-party ID to Snowflake's [`EXTERNAL_OAUTH_AUDIENCE_LIST`](https://docs.snowflake.com/en/sql-reference/sql/create-security-integration-oauth-external) for the security integration. * Confirm which claim your ADFS access token uses to carry the Snowflake username (for example, `email`), and set Snowflake's `EXTERNAL_OAUTH_TOKEN_USER_MAPPING_CLAIM` to match. Prophecy reads this same claim from the access token to determine the session user. If the claim is missing or does not match the Snowflake configuration, login fails with an error that identifies the claim mismatch. # AI settings Source: https://docs.prophecy.ai/administration/management/teams/settings/ai Configure AI settings for a team Applicable to the [Enterprise Edition](/administration/platform/editions) only. ## Accessing Copilot settings To configure Copilot settings for a team: 1. Open the **Metadata** icon in the left navigation bar. 2. Select the **Teams** tab in the page header. 3. Click on the team to open its metadata page. 4. Click the **Settings** tab. 5. Navigate to the **Advanced** subtab. Only team admins can access and modify AI settings for their team. ## Copilot settings The Copilot settings section contains two configuration options. ### Enable Copilot Toggle Copilot features (including Copilot and Agents) on or off for the team. * **Enabled**: Copilot features are available to all team members. * **Disabled**: Copilot features are not available to team members. ### Use data context in copilot When enabled, sample data from pipeline runs is used to improve Copilot suggestion quality. The Copilot analyzes actual data patterns and schemas from your runs to provide more relevant suggestions. * **Enabled**: Sample data from runs improves Copilot suggestion quality. * **Disabled**: Copilot suggestions are based on general patterns without data context. # Default dependencies Source: https://docs.prophecy.ai/administration/management/teams/settings/default-dependencies Configure default project dependencies for new projects Applicable to the [Enterprise Edition](/administration/platform/editions) only. Default projects are project dependencies that are automatically imported into new projects created within the team. This makes sense if your team frequently creates projects that use the same packages. Only team admins can access and modify default project dependencies for their team. ## Accessing default projects To configure default projects for a team: 1. Open the **Metadata** icon in the left navigation bar. 2. Select the **Teams** tab in the page header. 3. Click on the team to open its metadata page. 4. Click the **Settings** tab. 5. Navigate to the **Advanced** subtab. The Default Projects section displays all projects currently configured as defaults for the team. ## Default projects table The Default Projects table shows all projects that will be automatically imported into new projects. Each entry displays: * **Project**: The project name from Package Hub. * **Description**: Optional description of the project's purpose. * **Version**: The specific version of the project that will be imported. New projects will automatically import the latest version of the default projects. ## Managing default projects ### Adding default projects To add projects to the default list: 1. Open the **Add projects** dropdown. 2. Select projects from the Package Hub that you want to include as defaults. 3. The selected projects appear in the table. You can select any project from the Package Hub to include as a default dependency. If you add default dependencies written in different languages, Prophecy will only automatically import the dependencies written in the same language as the project. ### Removing default projects To remove a default project: 1. Locate the project in the Default Projects table. 2. Hover the project you want to remove. 3. Click the **trash** icon. # Project creation templates Source: https://docs.prophecy.ai/administration/management/teams/settings/project-creation-template Define a set of project creation settings Applicable to the [Enterprise Edition](/administration/platform/editions) only. At the team level, you can create multiple templates that include predefined project creation parameters. This is useful if your team creates many projects that use the same settings when they are created. **Only team admins can access the template settings for their team.** ## Template settings To find existing templates or add new templates: 1. Open the relevant **team** metadata. 2. Click on the **Settings** tab. 3. Navigate to the **Default Project Settings** subtab. Here, you can review or update existing templates by clicking through the **Templates** dropdown. To create a new template, click **+ Add New**. ## Parameters When you add a new template for your team, you need to fill in the following parameters. Parameters will differ between Spark and SQL projects. ### Spark | Parameter | Description | | -------------------------- | ---------------------------------------------------------------------------------------------------------- | | Template Name | A name to identify the template. | | Project Type | The project language. Select Python or Scala for Spark projects. | | Allowed to customize | Whether the template settings can be modified when a user creates a project with the template (Yes or No). | | Git Provider | Whether the project will use Prophecy-managed Git or an external Git provider. | | Git Storage Model | The option to choose **Normal** Git or **Fork per User**. | | Select as Default template | The option to make the template preselected during project creation. | | Default Main Branch | The default main branch name. | ### SQL | Parameter | Description | | -------------------------- | ---------------------------------------------------------------------------------------------------------- | | Template Name | A name to identify the template. | | Project Type | The project language. Select SQL for SQL projects. | | Allowed to customize | Whether the template settings can be modified when a user creates a project with the template (Yes or No). | | SQL Provider | The SQL provider that will execute dbt. | | Default Main Branch | The default main branch name. | | Default Development Branch | The default development branch name. This field appears **only if using external Git**. | | Select as Default template | The option to make the template preselected during project creation. | | Git Provider | Whether the project will use Prophecy-managed Git or an external Git provider. | | Git Storage Model | The option to choose **Simple** Git, **Normal** Git, or **Fork per User**. | ## Usage When you have templates for a certain team, users who select that team during project creation will see those templates. If you select a customizable template, you will still be able to make changes to settings during project creation. # Spark settings Source: https://docs.prophecy.ai/administration/management/teams/settings/spark Configure Spark settings for a team Available for [Enterprise Edition](/administration/platform/editions) only. The following sections describe different tabs within a team's metadata settings that pertain to **Spark projects**. ## Execution Metrics In the Execution Metrics tab, you can enable or disable [execution metrics](/data-engineering/fabrics/execution-metrics) at the team level. You can also specify in which tables to store execution metrics for each pipeline run. ## Code Generation When the **Enable multi file code generation** setting is enabled, Prophecy splits interactive execution code into multiple files instead of generating a single large file. Interactive execution code (generated when clicking the play button) differs from the project code committed to Git - it's optimized for running in Spark shell environments. This optimization reduces code recompilation by only updating changed portions of the pipeline during interactive runs, improving execution performance. ## Advanced In the Advanced tab, the following configuration is relevant to Spark projects: * **Artifact ID:** Sets the parent artifact identifier for project JAR and wheel packages. This identifier is used when packaging your Spark project dependencies and becomes part of the artifact metadata for deployment and distribution. # Teams Source: https://docs.prophecy.ai/administration/management/teams/teams Manage team structure and resources Teams are groups of users who collaborate on [projects](/data-analysis/development/projects/create-project) and share access to resources such as [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics). When you create a project or a fabric, you assign it to a team. All users in that team will have access to the relevant project or fabric. ## Team types There are two types of teams in Prophecy. * **Personal teams:** When you start using Prophecy, you are automatically assigned to your own one-person team. You are also the team admin of this team. If you want a project or fabric to be accessible only to yourself, assign it to your personal team and keep it private. * **Shared teams:** Your [team admin](/administration/management/users/access/role-based-access) typically creates additional team groupings. Team structures will vary across organizations. To learn about best practices for organizing teams, visit [Team-based access](/administration/management/users/access/team-based-access). ## Team admins Team admins are responsible for managing the structure and resources of a team. They can: * [Create and manage](/administration/management/users/team-user-provisioning) teams and team members. * Create and configure fabrics for the team. * Set up connections and secrets. * Manage additional team-level settings. There can be multiple team admins per team, and any existing team admin can promote another user to this role. For more information, visit [Role-based access control](/administration/management/users/access/role-based-access). Note that team admins are different from Prophecy cluster admins, who manage infrastructure, authentication, and deployment-wide settings. ## Team metadata Manage the entities within a team by accessing the team's metadata page. 1. Click the **Metadata** icon in the left navigation bar. 2. Select the **Teams** tab in the page header. 3. Click the team to open its metadata page. The team metadata page includes: * Entities that the team owns, such as projects, pipelines, datasets, and jobs. * Information on team members. * Settings that are only visible to team admins. # Credits Source: https://docs.prophecy.ai/administration/management/usage-billing/credits How credit consumption in Prophecy corresponds to platform usage Credits apply to the [Free and Professional Editions](/administration/platform/editions) only. In Prophecy, **credits** represent a unified measure of usage across the platform. Plans, teams, and fabrics define how billing and credit consumption are organized. * A plan is the top-level billing unit. Each plan includes exactly one team. * A [team](/administration/management/teams/teams) represents a group of users who share access to resources through one fabric. * A [fabric](/data-analysis/environment/fabrics/prophecy-fabrics) points to a dedicated Prophecy warehouse database and compute resources. To learn about monitoring credit usage and buying additional credits, see [Usage and billing](/administration/management/usage-billing/usage-billing). ## How credits are charged The following actions consume the corresponding number of credits. | Action | Description | Credits Consumed | | ----------------- | -------------------------------------------------------- | ---------------- | | **Team Members** | Number of users allowed on your plan | 20 / user | | **AI Requests** | Each time you click on Copilot or interact with an Agent | 0.04 / request | | **Compute** | Processing required when running a gem or pipeline | 3 / CPU hour | | **Data Egressed** | Data sent outside the Prophecy warehouse | 0.045 / GB | | **Data Storage** | Data stored in the Prophecy warehouse | 0.02 / GB | The Prophecy warehouse is powered by DuckDB. Advanced users can update the DuckDB SQL queries in the **Code** view of their project to optimize computation and reduce credit consumption. ## Relationship diagram The following diagram shows how credit consumption actions relate to plans, teams, and fabrics. ```mermaid theme={null} flowchart LR A[Plan] --> B[Team] B --> C[Fabric / Prophecy Warehouse] subgraph "Credit Consumption" direction LR E[Team Members] F[AI Requests] G[Compute] H[Data Storage] I[Data Egressed] end B -.-> E C -.-> F C -.-> G C -.-> H C -.-> I %% Styling classDef entity fill:#f5f5f5,stroke:#bdbdbd,stroke-width:1px,color:#1c1e21 classDef metric fill:#e1f5fe,stroke:#bdbdbd,stroke-width:1px,color:#1c1e21 classDef cluster fill:transparent,stroke:#bdbdbd,stroke-width:1px class A,B,C entity class E,F,G,H,I metric ``` ## What happens when you run out of credits? When you run out of credits, the following features will be disabled: * Prophecy Agent * Interactive pipeline runs * Scheduled pipeline runs * Project publication (deployment) # Manage billing Source: https://docs.prophecy.ai/administration/management/usage-billing/usage-billing Learn how billing works across Prophecy editions and how to manage usage and costs Prophecy billing is determined by edition. Self-serve billing is available for Free and Professional Editions. Express and Enterprise Editions are managed through external purchase channels. | Edition | Billing | | ------------ | ----------------- | | Free | Stripe | | Professional | Stripe | | Express | Cloud Marketplace | | Enterprise | Purchase Order | ## Free and Professional Editions Free and Professional Editions use a credit-based billing model managed through Stripe. * **Free Edition**: 5 credits per month. * **Professional Edition**: includes a fixed number of base credits per month; extra usage is pay-as-you-go. ### Plan-team relationship Each Prophecy plan is associated with a single team. [Team admins](/administration/management/teams/teams) manage billing for their team. Review the following example: 1. An existing team admin (User A) invites a new user (User B) to their team. 2. User B signs up for Prophecy and joins the existing team as a standard user. Team admins can later promote them to admin if needed. 3. User B automatically receives their own personal team under the Free Edition plan, where they are the team admin. They can choose to upgrade this team to the Professional Edition. This model ensures that users can have separate personal teams with independent plans. The flow is illustrated below: ```mermaid theme={null} graph TD %% Users and their teams UserA[User A] --> TeamA[Professional Edition Plan
Team A
2 users] UserA -->|invites| UserB UserB[User B] --> TeamA UserB --> TeamB[Free Edition Plan
Team B
1 user] %% Each team has a fabric TeamA -->|auto-provisions| FabricA[Fabric A] TeamB -->|auto-provisions| FabricB[Fabric B] %% Styling classDef user fill:#e1f5fe,stroke:#bdbdbd,stroke-width:1px,color:#1c1e21 classDef team fill:#f3e5f5,stroke:#bdbdbd,stroke-width:1px,color:#1c1e21 classDef other fill:#f5f5f5,stroke:#bdbdbd,stroke-width:1px,color:#1c1e21 class UserA,UserB user class TeamA,TeamB team class FabricA,FabricB other ``` ### Usage & Billing dashboard The **Usage & Billing** dashboard displays plan information, usage, and credit consumption. Because each plan maps to an individual team, select a specific team to view its corresponding billing information. To access the dashboard: 1. Log in to Prophecy and open the **Settings** page. 2. Select the **Usage & Billing** tab. 3. Use the team toggle to view usage for different teams. From the dashboard, team administrators can: * Monitor credit usage. * Set seat limits. * Configure alerts and budgets. * Update payment method. For information about how credits are consumed, see [Credits](/administration/management/usage-billing/credits). ### Manage monthly spending To set up monthly spending alerts and limits: 1. Go to **Settings β†’ Usage & Billing**. 2. Under **Additional Usage**, click **Manage**. Alternatively, click the **Manage monthly spending** button on the page. * **Set a user seat limit**: Maximum number of users allowed for the team corresponding to the plan. * **Set a usage alert**: Usage amount in dollars that triggers an alert and email. * **Set a usage budget**: Usage amount in dollars at which services are suspended. 3. Change the values and click **Save**. Review the default values when you first upgrade to the Professional Edition, and update if needed. ### Update payment method To update your payment method: 1. Go to **Settings β†’ Usage & Billing**. 2. Under **Payment method**, click **Manage**. 3. Follow the instructions in Stripe to update your payment method. ### Upgrade plan To upgrade from the Free Edition to the Professional Edition: 1. Go to **Settings β†’ Usage & Billing**. 2. Selected the correct team to upgrade in the team toggle. 3. Scroll to the bottom of the page to view the different editions. 4. In the **Pro** tile, select **Upgrade Now**. You will be taken to Stripe for payment. ### Downgrade plan To downgrade from Professional Edition to the Free Edition: 1. Go to **Settings β†’ Usage & Billing**. 2. Under **Free Edition**, select **Request Downgrade**. This will initiate steps to cancel your Professional Edition plan. ## Express Edition Express Edition is billed through the cloud marketplace (AWS or Azure) where the deployment was initialized. Billing and invoicing are handled by the marketplace provider, and details are available through their respective dashboards. ## Enterprise Edition Enterprise Edition is billed through purchase orders. Invoicing, contracts, and terms are managed directly between Prophecy and your organization. Billing information is not available in the Prophecy UI. To upgrade to the Enterprise Edition, you must [contact us](https://www.prophecy.io/contact-us). You won't be able to upgrade independently. # Role-based access control (RBAC) Source: https://docs.prophecy.ai/administration/management/users/access/role-based-access Manage access by role Prophecy uses role-based access control (RBAC) to manage permissions across users and teams. Each user is assigned one or more roles that determine what actions they can perform and what resources they can access. There are three key roles in Prophecy: * **Standard users**: Can access and work within the teams they belong to * **Team admins**: Manage team membership and team-level resources like fabrics and connections * **Prophecy cluster admins ([Enterprise](/administration/platform/editions) only)**: Administer the overall Prophecy deployment and infrastructure All users start as standard users by default, with the exception of being their [personal team](#personal-team) admin. They can be granted the team admin role for specific teams, and automatically become the team admin for any team they create. In contrast, Prophecy cluster admin roles are managed by Prophecy and assigned at the deployment level. ## Standard users Standard users access resources through team assignments. In other words, permissions are governed by team-level access controls. Standard users can: * Create projects within their teams * Attach to fabrics that are assigned to their teams Users cannot create or edit fabrics for a team unless they are also a [team admin](#team-admins). For more information about best practices, visit [Team-based access](/administration/management/users/access/team-based-access). ### Personal team Every user is automatically given a personal team, named after their login email. This team includes only the user and grants them team admin permissions, allowing them to create both projects and fabrics for individual use. ## Team admins Team admins manage teams and create resources for their [teams](/administration/management/teams/teams). This includes responsibilities like: * Adding and removing users from teams * Creating fabrics that correspond to different execution environments * Setting up connections with the appropriate credentials * Deploying projects to run scheduled pipelines * **([Professional Edition](/administration/platform/editions) only)** Managing usage and billing for the team The user who creates a team is automatically assigned as its team admin. Additional team admins can be added or disabled for each team. For recommendations regarding team setup and organization, visit [Team-based access](/administration/management/users/access/team-based-access). ## Prophecy cluster admins You can assign this role to users in the Enterprise Edition only. Prophecy cluster admins manage clusters, infrastructure, compute resources, and Prophecy deployment. This includes responsibilities such as: * Setting up authentication like SSO for the Prophecy environment * Managing audit log review and storage * Upgrading Prophecy to a newer version * Downloading system logs to send to Prophecy's Support team Prophecy automatically provisions one Prophecy cluster admin per Enterprise Edition deployment. Additional cluster admins can be created if required. # Team-based access Source: https://docs.prophecy.ai/administration/management/users/access/team-based-access Manage access by team In Prophecy, access to resources is based on team associations with resources. Projects and fabrics are the fundamental resources to which teams are assigned. Users who belong to a team can access all projects and fabrics assigned to that team. Relationship between users, teams, projects, and
fabrics This guide provides comprehensive best practices for implementing and managing team-based access to optimize collaboration while maintaining appropriate security boundaries. [Team admins](/administration/management/users/access/role-based-access) have special privileges for managing team membership and team fabric creation, but the core access model remains focused on the team-to-resource relationship. Each team can be assigned to many projects and fabrics, but importantly, each project and fabric can only be assigned to one team. ## Team structure The way you structure your teams will fundamentally shape how users interact with data and resources in Prophecy. There are several approaches to team organization, each with its own advantages depending on your organization's needs. In any case, teams should be assigned only to the resources they genuinely need to access. This principle minimizes the potential impact of compromised credentials and reduces the risk of accidental data modifications. ### Strategy 1: Environment-based teams One of the most straightforward approaches is to split teams based on access to development and production environments. This creates a clear separation between experimentation and production workloads. #### Dev Team (All Users) * Limited access to data * Access to development environment(s) only * Allows for safe experimentation and testing #### Prod Team (Platform Team Only) * Complete access and control to production environment * Restricted to users responsible for production operations * Ensures controlled deployment to production ### Strategy 2: Function-based teams As organizations scale, they often benefit from more granular team structures aligned with specific job functions. This approach allows for more precise resource allocation and specialized access patterns tailored to different roles within the data ecosystem. Here we provide one example of function-based teams. #### Data Engineering Team * Access to all environments, especially with raw data * Design and build data pipelines * Manage data quality and optimization #### Marketing Analytics Team * Access to environment with web analytics and campaign data * Create reporting pipelines * Send reports to BI tools to create dashboards * Run [analyses](/data-analysis/analysis/overview) #### Product Analytics Team * Access to environment with user behavior data and product metrics * Create reporting pipelines * Send reports to BI tools to create dashboards * Run [analyses](/data-analysis/analysis/overview) ## Fabric creation Fabrics are the bridge between Prophecy and your execution environment. They encapsulate connection details and access credentials, making them a critical component of your security architecture. Thoughtfully configured fabrics ensure that the right people have the right level of access to your data processing resources. Authentication capabilities vary between different connection types across fabrics. Visit individual connection documentation to learn about the specific authentication strategies available for that connection. ### Development fabrics Development fabrics should be configured to facilitate collaborative work and allow for experimentation. * **Team assignment**: Dev team * **Connection authentication**: Individual authentication * **Benefits**: Maintains accountability with each run linked to an individual identity * **Usage**: For projects in development phase only Here is one example of how you may configure a [Prophecy fabric](/data-analysis/environment/fabrics/prophecy-fabrics) for development. Development fabric ### Production fabrics Production fabrics require a different approach focused on stability and reliability. These environments run your business-critical processes and must operate consistently without frequent intervention. * **Team assignment**: Prod team * **Authentication**: Service principal * **Benefits**: Eliminates reauthentication needs and ensures smooth job execution * **Usage**: For scheduled pipelines in deployed projects Using service principals for production fabrics ensures that scheduled jobs continue to run even when individual team members are unavailable or their credentials change. Here is one example of how you may configure a [Prophecy fabric](/data-analysis/environment/fabrics/prophecy-fabrics) for production. Production fabric ## Project management Projects in Prophecy encapsulate data pipelines, gems, tables, and other components. How you structure and assign these projects to teams will differ between use cases. When you first begin using Prophecy, you can create projects with personal team for experimentation, then migrate to collaborative projects organized by business objectives: * **Create purpose-specific projects**: "Marketing Analytics," "Customer Retention," etc. * **Align with team structure**: Ensure projects are assigned to appropriate teams * **Document project scope**: Define project boundaries and ownership Purpose-specific projects help maintain focus and prevent the sprawl of unrelated components within a single project. By clearly defining each project's boundaries, teams can more easily understand what belongs where and how different components relate to business objectives. ### Sharing and collaboration Prophecy's sharing model enables teams to reuse project components without compromising governance. Projects can be shared with other teams to extend access. When a project is shared with your team: * You can import project components into your own projects through the [Package Hub](/data-engineering/extensibility/package-hub/package-hub). * You can run [analyses](/data-analysis/analysis/overview) built on top of project pipelines. * You **cannot** modify original project components. This model encourages the creation of reusable components that can be leveraged across teams without duplication. When a team creates a high-quality transformation or dataset definition, other teams can incorporate it into their workflows without risk of breaking the original implementation. ## Project and fabric relationships The relationship between projects and fabrics is a fundamental aspect of Prophecy's architecture that directly impacts your team-based access strategy. Let's see how these elements interact. ### Execution requirements Prophecy projects must be connected to a fabric to run pipelines. This is because fabrics contain your connections to execution environments and data sources that power your pipelines. Without this connection, a project remains a blueprint that cannot be executed. ### Cross-team connections A key aspect of this relationship is that project and fabric team assignments operate independently. You can connect a project assigned to one team to a fabric assigned to a different team, provided you are a member of both teams. This independence lets you perform actions such as promoting projects from development to production environments. For example, a business analyst in the development team can develop and test a project using a development fabric. When ready for production, a platform engineer in both the development team and production team can attach a production fabric to the project. Attach different fabrics to a
project To accomplish the scenario above, you must ensure that at least one user is a member of both the project's team and the target fabric's team. This way, someone will be able to publish or deploy a project to production environments. ### Provider compatibility While cross-team connections provide flexibility, there are technical constraints to consider. For SQL projects **only**, the project must have the same provider as the fabric. This means: * Databricks projects can only attach to Databricks fabrics. * Snowflake projects can only connect to Snowflake fabrics. Python and Scala projects do not specify "providers" at creation. These projects can connect to any Spark fabric. # Account settings Source: https://docs.prophecy.ai/administration/management/users/account-settings Update your Prophecy account details You can update your account details in the **Info** tab of the **Settings** interface in Prophecy. ## Update account details Use the following steps to update your first name, last name, or company associated with your Prophecy account. 1. Navigate to **Settings β†’ Info**. 2. Hover over the field you want to change. 3. Click the **pencil** icon to edit. 4. Type in the new information and click **Enter**. 5. Click **Update** to save your changes. You cannot change the email address of your Prophecy account. ## Change password To change your account password: 1. Navigate to **Settings β†’ Info**. 2. Under **Change Password**, fill out the required fields. 3. At the bottom of the page, click **Change Password**. # Team and user provisioning Source: https://docs.prophecy.ai/administration/management/users/team-user-provisioning Learn how to create new teams and invite users This page explains how to add users and teams to your Prophecy environment. For guidance on how to configure your teams, visit [Team-based access](/administration/management/users/access/team-based-access). ## Create teams Before you can invite users to Prophecy, you need to define the [teams](/administration/management/teams/teams) they will be invited to. 1. Go to **Metadata β†’ Teams**. 2. Click **+ Create Team**. 3. In the **Basic Info** tab, define the team name. 4. Click **Continue**. 5. **([Enterprise Edition](/administration/platform/editions) only)** Enable or disable Spark execution metrics for the team. 6. Click **Complete** to save the new team. ## Create users After you define the teams, invite new users to those teams. You can do so in one of two ways: * Directly through the **Metadata** page. * By editing a team's settings ### Add team members through the Metadata page To add team members through the Metadata page: 1. Go to **Metadata β†’ Users**. 2. Click **+ Invite Users**. 3. In the **Invite User** dialog: * Enter one or more email addresses. * Select the team the user will join. * Choose a role: **User** or **Admin**. * Click **Send Invitation**. The system creates the accounts and emails the invitation links to users. ### Add members through the Teams tab To add team members by editing a team: 1. Go to **Metadata β†’ Teams**. 2. Click the link for your team. 3. Click the **Members** tab. 4. Click **+ Invite Users**. 5. In the **Invite User** dialog: * Enter one or more email addresses. * Select the team the user will join. * Choose a role: **User** or **Admin**. * Click **Send Invitation**. The system creates the accounts and emails the invitation links to users. # Architecture Source: https://docs.prophecy.ai/administration/platform/architecture Understand the infrastructure behind a Prophecy deployment Prophecy operates as a distributed system built on microservices architecture, orchestrated by Kubernetes across multiple cloud platforms. The platform consists of several core components that work together to provide data transformation, orchestration, and management capabilities. ## Architecture by Edition ### Free and Professional Edition The Free and Professional Editions provide a complete data platform with managed components. | Component | Description | | ------------------ | -------------------------------------------------------------------------------------------------------------- | | Prophecy Studio | The control plane that provides the user interface for developing visual data pipelines and managing projects. | | Prophecy Automate | The native runtime designed for data ingestion, egress, and built-in scheduling capabilities. | | Prophecy In Memory | The Prophecy-managed SQL warehouse that processes data transformations. | | Data storage | Data outside of the execution environment that will flow in and out of the pipeline. | | AI endpoint | Prophecy-managed LLM subscription and endpoint. | | Version control | Git integration supporting both Prophecy-managed and external Git repositories. | | Deployment model | SaaS only. Learn more in [Deployment models](/administration/platform/prophecy-deployment). | Free and Professional Edition Architecture ### Express Edition The Express Edition provides enterprise-grade features scoped to leverage your existing SQL warehouse infrastructure. | Component | Description | | ---------------------- | -------------------------------------------------------------------------------------------------------------- | | Prophecy Studio | The control plane that provides the user interface for developing visual data pipelines and managing projects. | | Prophecy Automate | The native runtime designed for data ingestion, egress, and built-in scheduling capabilities. | | External SQL Warehouse | Your own Databricks SQL engine that executes data transformations. | | Data storage | Data outside of the execution environment that will flow in and out of the pipeline. | | AI endpoint | Customer-managed LLM subscription and endpoint. | | Version control | Git integration supporting both Prophecy-managed and external Git repositories. | | Deployment model | Dedicated SaaS only. Learn more in [Deployment models](/administration/platform/prophecy-deployment). | Enterprise and Express Edition Architecture This diagram shows the architecture for the Express Edition. **Users on the Enterprise Edition can also leverage this architecture.** However, Enterprise users can also connect to additional SQL warehouses, like BigQuery and Snowflake. ### Enterprise Edition The Enterprise edition offers maximum flexibility with multiple execution engine options and deployment models. | Component | Description | | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Prophecy Studio | The control plane that provides the user interface for developing visual data pipelines and managing projects across various data platforms. | | Execution engine | Flexible compute options including Spark clusters or external SQL warehouses combined with Prophecy Automate. Prophecy executes data transformations on your chosen execution environment. Fabrics enable users to execute pipelines on these platforms. Prophecy does not persist your data. | | Data storage | Data outside of the execution environment that will flow in and out of the pipeline. | | AI | Customer-managed LLM subscription and endpoint. | | Version control | Git integration supporting both Prophecy-managed and external Git repositories. | | Deployment model | Dedicated SaaS preferred and SaaS available. Learn more in [Deployment models](/administration/platform/prophecy-deployment). | Enterprise Edition Spark Architecture The Enterprise Edition supports both SQL-based and Spark-based architectures. The diagram above shows the architecture for a deployment using Spark. Prophecy can accommodate a wide variety of architectures beyond this diagram. For example: * The diagram demonstrates Databricks as the execution engine. You can connect to other platforms like Amazon EMR and Google Cloud Dataproc, or use another Spark engine through [Apache Livy](https://livy.apache.org/). * The diagram displays a connection to an external Git repository. You can connect to a variety of providers such as GitHub, Bitbucket, GitLab, and more. ## What is Prophecy Automate? Prophecy Automate is the native runtime available across all Prophecy editions. It extends pipeline execution by providing orchestration, external data handling, and observability for both SQL and PySpark pipelines. * **Ingress / Egress**: Automate enables pipelines to interact with external systems such as databases, file storage, and APIs. * **Orchestration**: Automate manages pipeline execution across environments through time-based and trigger-based scheduling, as well as orchestration APIs for programmatic control. It also supports dynamic execution patterns through features like the [DynamicInput gem](/data-analysis/gems/custom/dynamic-input), which adapts queries based on incoming data, and the [Directory gem](/data-analysis//gems/custom/directory), which enumerates files and folders from storage systems. * **Observe**: Automate provides visibility into pipeline activity, including run history and status tracking, project deployment tracking, and monitoring of active pipeline schedules. Prophecy Automate connections diagram ## Supported compute engines | Provider | SQL Pipelines | Spark Pipelines | SQL Models | | ------------------ | ------------- | --------------- | ---------- | | Prophecy In Memory | βœ” | | βœ” | | Databricks | βœ” | βœ” | βœ” | | BigQuery | βœ” | | βœ” | | Snowflake | βœ” | | βœ” | | EMR | | βœ” | | | Dataproc | | βœ” | | | Synapse | | βœ” | | | Livy | | βœ” | | # Editions Source: https://docs.prophecy.ai/administration/platform/editions Learn about and compare the different Prophecy editions Prophecy is available in multiple editions to support a range of use cases. The Free and Professional Editions allow you to get started quickly, with Prophecy managing infrastructure and resources. The Express Edition is designed for organizations already using Databricks SQL, Google BigQuery, or Snowflake. The Enterprise Edition provides the most flexibility and control, with advanced options for deployment, security, and platform management. The sections below outline differences across deployment, security, compute, AI, platform, sharing, and support to help you evaluate which edition best fits your requirements. ## Free and Professional Edition Free and Professional Editions are designed for data analyst teams that want to collaborate in real time without managing infrastructure. Both editions provide Prophecy-managed resources that are metered by credits. The Free Edition is functionally identical to the Professional Edition, but is limited to 5 credits per month and a single user per plan. First, sign up for the Free Edition. You'll be able to upgrade to the Professional Edition after you sign in for the first time. ## Express Edition Express Edition is designed for organizations that already use Databricks SQL, Google BigQuery, or Snowflake and want to extend existing infrastructure with Prophecy. Deployments are provisioned as Dedicated SaaS environments through the AWS, Azure, or Google Cloud marketplaces. It is the right choice for teams standardizing on Databricks, Snowflake, or BigQuery who also need options for platform customization and integration with enterprise networking features such as PrivateLink. Express Edition supports up to 20 users, with no more than 5 simultaneous users recommended. * Get started on [Azure Marketplace](https://azuremarketplace.microsoft.com/en-us/marketplace/apps/simpledatalabsinc1635791235920.prophecy-enterprise-express-for-databricks?tab=Overview). * Get started on [AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-dht7vktn2yues). * Get Started on [Google Cloud Marketplace](https://console.cloud.google.com/marketplace/product/prophecy-on-gcp-public/prophecy-enterprise-express). ## Enterprise Edition Enterprise Edition is designed for organizations that require the highest level of control, security, and flexibility. Deployments are hosted in a dedicated Prophecy VPC, with support for advanced configuration across networking, compute, and platform services. Enterprise Edition is best suited for organizations that need integration with existing enterprise systems, including support for SCIM, single sign-on, PrivateLink, and bring-your-own-key encryption. It also provides access to audit logs and other compliance features required in regulated industries. To get started with the Enterprise Edition, [reach out to Prophecy](mailto:contact.us@prophecy.io). To try out Enterprise features, sign up for our [multi-tenant SaaS Enterprise environment](https://app.prophecy.io/). ## Feature matrix | Feature | Free | Professional | Express | Enterprise | | ---------------------------------------------------------------------------------------------------------- | ---- | ------------ | ------- | ---------- | | **Deployment** | | | | | | SaaS (Multi-tenant) | βœ” | βœ” | | | | Dedicated SaaS (Single-tenant) | | | βœ” | βœ” | | Feature flag customization | | | βœ” | βœ” | | Maximum user limits | βœ” | | βœ” | | | Usage metered by credits | βœ” | βœ” | | | | **Security** | | | | | | IP Whitelisting | βœ” | βœ” | βœ” | βœ” | | PrivateLink | | | βœ” | βœ” | | Google SSO | βœ” | βœ” | βœ” | βœ” | | Additional [authentication](/administration/management/authentication/authentication) providers | | | βœ” | βœ” | | Automatic team creation using [SCIM](/administration/management/authentication/saml) | | | | βœ” | | Bring Your Own Key (BYOK) | | | | βœ” | | Audit log access | | | | βœ” | | **Compute** | | | | | | [Prophecy Automate](/administration/platform/architecture) | βœ” | βœ” | βœ” | βœ” | | [Prophecy In Memory](/data-analysis/environment/fabrics/prophecy-fabrics) (Prophecy-managed SQL warehouse) | βœ” | βœ” | | | | Self-managed SQL warehouse | | | βœ” | βœ” | | Self-managed Spark | | | | βœ” | | Automatic fabric provisioning | βœ” | βœ” | | | | Manual fabric setup | | | βœ” | βœ” | | **AI** | | | | | | Agents | βœ” | βœ” | βœ” | βœ” | | Copilot | βœ” | βœ” | βœ” | βœ” | | Prophecy-managed LLM subscription | βœ” | βœ” | | | | Self-managed LLM subscription | | | βœ” | βœ” | | **Platform features** | | | | | | Git-hosted projects | βœ” | βœ” | βœ” | βœ” | | Build custom gems | βœ” | βœ” | βœ” | βœ” | | Analysis dashboards | βœ” | βœ” | βœ” | βœ” | | Pipeline run history and monitoring | βœ” | βœ” | βœ” | βœ” | | Prophecy-orchestrated pipelines | βœ” | βœ” | βœ” | βœ” | | Prophecy orchestration (Automate) | βœ” | βœ” | βœ” | βœ” | | External orchestration (Databricks Jobs, Airflow) | | | | βœ” | | Spark pipelines | | | | βœ” | | Models (dbt) | | | | βœ” | | Transpiler | | | βœ” | βœ” | | **Project sharing** | | | | | | Share by adding users to the project team | βœ” | βœ” | βœ” | βœ” | | Share via public link | βœ” | βœ” | | | | Share project replays | βœ” | βœ” | | | | Publish to Package Hub | | | βœ” | βœ” | | **Support** | | | | | | Community forums | βœ” | βœ” | βœ” | βœ” | | Support with SLAs | | | | βœ” | Usage of the Prophecy-managed SQL warehouse is metered by credits. # Deployment models Source: https://docs.prophecy.ai/administration/platform/prophecy-deployment Understand the SaaS and Dedicated SaaS deployment models Every Prophecy edition runs on one of two deployment models: SaaS or Dedicated SaaS. Which model you're on depends on your edition. This page explains what each model means for tenancy, isolation, and who's responsible for what. | Feature | SaaS | Dedicated SaaS | | ---------------------------- | ---- | -------------- | | Prophecy-managed upgrades | βœ” | βœ” | | Prophecy-managed maintenance | βœ” | βœ” | | Multi-tenancy | βœ” | | | Single-tenancy | | βœ” | | Customizable environment | | βœ” | For a breakdown of which components and compute engines are available on each edition, see [Architecture](/administration/platform/architecture). ### SaaS SaaS is entirely Prophecy-managed and runs on a multi-tenant architecture in Prophecy's AWS VPC (virtual private cloud), with multiple layers of logical isolation between tenants. This deployment model gives you the fastest access to new features and updates. Sign up for a [free trial](https://app.prophecy.io/metadata/auth/signup) to evaluate Prophecy on SaaS. ### Dedicated SaaS Dedicated SaaS requires the [Express or Enterprise Edition](/administration/platform/editions) of Prophecy. Express Edition runs on Dedicated SaaS only; Enterprise Edition can run on either SaaS or Dedicated SaaS. Dedicated SaaS combines Prophecy-managed infrastructure with the privacy and isolation of a single-tenant architecture, deployed to a dedicated Prophecy VPC on AWS, Google Cloud, or Azure. This runs in the region of your choice in order to support data residency requirements. Prophecy manages installation, maintenance, and resource allocation for you. Prophecy connects outbound over TLS to your identity provider, Git repository, data platform, and (if configured) your AI LLM endpoint/gateway. For details on what runs in each plane, see [Architecture](/administration/platform/architecture). #### Data storage and security Your data stays in your own SQL warehouse and Prophecy never stores or copies it. When you preview a dataset, sample data passes through Prophecy's execution service only momentarily to render it in your browser; data is no persisted anywhere in Prophecy itself. Prophecy does store metadata: schema, object names, relationships, and other configuration information, at a level you can control. Prophecy encrypts this information both at rest (using AES-256, via your cloud provider's key management service) and in transit (TLS 1.3). Dedicated SaaS supports Bring Your Own Key (BYOK) for encryption. #### Responsibility matrix This table outlines the division of responsibilities between customers and Prophecy for Dedicated SaaS deployments. | Area | Customer Responsibility | Prophecy Responsibility | Description | | --------------------------------------- | :---------------------: | :---------------------: | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Platform upgrades | | βœ” | Prophecy applies upgrades and hotfixes. | | Security and compliance | | βœ” | Prophecy applies CVE patches, manages SaaS compliance posture, and ensures tenant isolation. | | High availability and disaster recovery | | βœ” | Prophecy implements failover strategies and backup/restore procedures, and tests disaster recovery scenarios to maintain system resilience. | | Kubernetes cluster and infrastructure | | βœ” | Prophecy manages scaling, monitoring, logging, namespaces, and storage for both the Control Plane and Execution Plane. | | Scaling and resource tuning | | βœ” | Prophecy optimizes performance, adjusts resource allocation, and configures auto-scaling. | | Identity and access management | βœ” | | Customer configures users and groups in their chosen IdP, then sets up SSO inside Prophecy. | | Networking | βœ” | βœ” | Prophecy manages a list of up to 300 IP addresses allowed to access the Dedicated SaaS deployment, connects outbound from a single IP address, and maintains networking infrastructure. Customer accepts connection requests and allowlists Prophecy's IP address with Databricks, GitHub, and other external platforms Prophecy needs to reach. | | Data encryption (Bring Your Own Key) | βœ” | βœ” | Customer optionally provides and manages a Key Management Service (KMS) and grants Prophecy access to customer-managed encryption keys. Prophecy integrates with customer KMS to encrypt persistent storage using those keys. | | Monitoring and logs | βœ” | βœ” | Customer reviews audit logs, optionally synced to a customer-owned storage bucket. Prophecy monitors SaaS infrastructure, performs root-cause analysis, updates the status page, and generates audit logs. Prophecy does not log metadata, data, credentials, secrets, or IP addresses (aside from browser IP addresses via the ingress controller). | | AI / LLM connectivity | βœ” (optional) | βœ” | By default, Prophecy connects to its managed foundation model provider. Customers can instead configure their own AI LLM endpoint or gateway. Prompts are encrypted in transit; Prophecy does not use customer data to train or fine-tune any model. | # Databricks Partner Connect Source: https://docs.prophecy.ai/administration/setup/databricks-partner-connect Get started with Prophecy via Databricks Partner Connect If you have access to a Databricks workspace, you can get started with Prophecy via [Databricks Partner Connect](https://docs.databricks.com/aws/en/partner-connect/). This integration lets you setup a Prophecy [Enterprise Edition](/administration/platform/editions) SaaS account directly from Databricks. ## Open the Prophecy integration First, you'll find the Prophecy integration within Databricks. 1. Open your Databricks workspace. 2. From the left sidebar, select **Marketplace**. 3. Under **Partner Connect integrations**, click **Prophecy**. 4. On the Prophecy Integration page, click **Connect**. In the dialog that pops up, you will configure the Databricks execution environment to use in Prophecy. 1. Choose the **SQL warehouse** that you'll use for pipeline execution in Prophecy. 2. Select the catalog and schema that Prophecy will use as the default for reading and writing data. You will only be able to select what is available to your Databricks user. 3. Click **Next > Next**. ## Sign up for Prophecy Prophecy's SaaS environment login should open in a new tab. 1. Enter a new password to create a Prophecy account. If you already have an account, sign in. 2. Open the **Metadata** page to see the list of projects you can access. 3. Open the HelloWorld\_SQL project in the project editor. 4. In the top right corner, attach the to SQL warehouse that you configured during the Partner Connect setup. Now that you have connected to your Databricks SQL warehouse, you can run pipelines and models in that environment. # Enterprise Edition Source: https://docs.prophecy.ai/administration/setup/prophecy-enterprise-guide Complete setup guide for Prophecy Enterprise Edition Prophecy [Enterprise Edition](/administration/platform/editions) is designed for organizations that need robust security, scalability, and customization. It provides advanced features such as SCIM-based user management, additional compute options, and secure networking configuration. **Dedicated SaaS** is the preferred deployment model for the Enterprise Edition. In this model, Prophecy runs in a secure, Prophecy-managed VPC that is isolated for your organization. This page describes how to configure a Dedicated SaaS deployment. ## Deployment setup The following are optional configurations for your deployment that you can work with Prophecy to set up. Please reach out to Prophecy if you want to enable any of the following configurations. * **SSO**: Enable single sign-on for Prophecy users. * **SCIM**: Auto-provision users into teams using SCIM groups. * **LLM**: Enterprise Edition requires a customer-provided LLM endpoint to power Prophecy AI. * **Transpiler**: Enable the Transpiler to migrate from external tools. * **Compute options**: * **SQL warehouse**: Databricks is enabled by default. You can also enable BigQuery or Snowflake. * **Prophecy Automate**: Prophecy Automate is enabled by default. You can choose to disable it. * **Serverless**: Option to connect to Databricks Serverless. ## User management Enterprise Edition allows an unlimited number of users in a Prophecy environment. Learn how to add teams and users in [Team and user provisioning](/administration/management/users/team-user-provisioning). ## Git credentials Prophecy projects are stored as code in Git repositories. You can use either: * A Prophecy-managed Git repository * An external Git repository To access external repositories, the corresponding Git credentials must be stored in Prophecy. * Learn how to [add new Git credentials](/data-engineering/ci-cd/git/git#Git-credentials). * Learn how to [share Git credentials](/data-engineering/ci-cd/git/git#share-credentials) with your team. When creating a new project, users can select to connect to the Prophecy-managed Git repository or any Git account they have access to in Prophecy. If a user selects a Git account, they can select any empty repository to store the project. ## Fabric configuration A fabric in Prophecy defines the execution environment for your pipelines. The Enterprise Edition supports a multiple fabric types. * [Prophecy fabrics](/data-analysis/environment/fabrics/prophecy-fabrics): Use Prophecy Automate with an external SQL warehouse to run pipelines. * [Spark fabrics](/data-engineering/fabrics/spark-provider/databricks/databricks): Use external Spark clusters to run pipelines. * SQL fabrics: Use an external SQL warehouse without Prophecy Automate. You can run models in the warehouse, but not pipelines. You can create individual fabrics for different execution environments (for example, `dev` and `prod`). ### Optional: Networking Dedicated SaaS deployments run in Prophecy's VPC. You may need to configure networking to allow Prophecy to communicate with your external services. Different connection types may have different networking requirements. If a connection fails, check that your networking configuration allows traffic between Prophecy and your external services. # Express Edition Source: https://docs.prophecy.ai/administration/setup/prophecy-express-guide Complete setup guide for Prophecy Express Edition Prophecy [Express Edition](/administration/platform/editions) is a streamlined version of Prophecy designed for business analysts who need to build and deploy data pipelines quickly. This edition provides essential data transformation capabilities with simplified administration and deployment options. ## LLM backend configuration Prophecy Express Edition requires a customer-provided endpoint for an LLM to power Prophecy AI. Contact Support to configure the model provider and LLM you will use, or request to disable AI features for your environment. ## Users and sign-on Express Edition supports up to 20 users, with no more than 5 simultaneous users recommended. Learn how to add teams and users in [Team and user provisioning](/administration/management/users/team-user-provisioning). ### Optional: Authentication Express Edition supports Single Sign-On (SSO) integration with Okta. Learn more about setting up Okta in our [SAML documentation](/administration/management/authentication/saml#set-up-saml). If you're not using SSO, users must use the exact same email addresses for their Prophecy accounts as they use for authenticating with Databricks. This ensures proper authentication flow between platforms. Express Edition **does not** support SCIM. ## Git credentials Prophecy projects are stored as code in Git repositories. You can use either: * A Prophecy-managed Git repository * An external Git repository To access external repositories, the corresponding Git credentials must be stored in Prophecy. * Learn how to [add new Git credentials](/data-engineering/ci-cd/git/git#Git-credentials). * Learn how to [share Git credentials](/data-engineering/ci-cd/git/git#share-credentials) with your team. When creating a new project, users can select to connect to the Prophecy-managed Git repository or any Git account they have access to in Prophecy. If a user selects a Git account, they can select any empty repository to store the project. ## Fabric configuration A fabric in Prophecy defines the execution environment for your pipelines. To set up an execution environment: * Create a [Prophecy fabric](/data-analysis/environment/fabrics/prophecy-fabrics). You can only create fabrics for teams where you are the [team admin](/administration/management/users/access/role-based-access). * To add SQL warehouse connection in the fabric, set up a [Databricks connection](/data-analysis/environment/connections/databricks). * To add data ingress/egress connections in the fabric, [follow the instructions](/data-analysis/environment/connections/connections) for each type of connection. You can create individual fabrics for different execution environments (for example, `dev` and `prod`). ### Optional: Networking Dedicated SaaS deployments run in Prophecy's VPC. You may need to configure networking to allow Prophecy to communicate with your external services. For example, to connect to Databricks, you might need to configure **Private Link** or set up **IP allowlists**. To do so: 1. Follow the instructions to set up PrivateLink depending on the cloud platform where Databricks is hosted. During this process, Databricks creates VPC endpoints for the secure cluster connectivity relay and for the workspace. * [Enable private connectivity using AWS PrivateLink](https://docs.databricks.com/aws/en/security/network/classic/privatelink) * [Enable Azure Private Link back-end and front-end connections](https://learn.microsoft.com/en-us/azure/databricks/security/network/classic/private-link) 2. Share the following with Prophecy: * Your Databricks workspace region. * The VPC endpoint service names. Provide both the relay and workspace endpoints. 3. Prophecy creates VPC endpoints in the Prophecy VPC that connect to those service names. 4. Prophecy shares the Prophecy VPC endpoints with you. 5. Whitelist the Prophecy endpoints in your Databricks configuration. Different connection types may have different networking requirements. If a connection fails, check that your networking configuration allows traffic between Prophecy and your external services. ## Workflow migration If you're migrating from Alteryx, Prophecy Express Edition includes a Transpiler to convert your existing workflows. 1. Use the **Alteryx Transpiler** to import your existing workflows. 2. Review and refine the converted pipelines. 3. Test the pipelines in the Prophecy environment. 4. Deploy the modernized workflows. For more information, visit our [Transpiler documentation](https://transpiler.docs.prophecy.io/). # Add connection to fabric Source: https://docs.prophecy.ai/api-reference/connections/add-connection-to-fabric post /api/orchestration/fabric/{fabricId}/connection Create a new connection in a specific fabric The `properties` object in the connection request body depends on the connection type. Find specific properties in the **Connections > Properties** section of the API documentation. # Delete connection Source: https://docs.prophecy.ai/api-reference/connections/delete-connection delete /api/orchestration/fabric/{fabricId}/connection/name/{connectionName} Delete a specific connection by its name # List connections per fabric Source: https://docs.prophecy.ai/api-reference/connections/list-connections-per-fabric get /api/orchestration/fabric/{fabricId}/connection List all connections and their details from a specific fabric # Azure Data Lake Storage connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/adls Properties required to create and update ADLS connections Properties for Azure Data Lake Storage (ADLS) connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties Azure Active Directory tenant ID. Azure Active Directory application (client) ID. Reference to a Prophecy secret containing the Azure client secret. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for client secrets. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Azure Data Lake Storage account name. Name of the container within the ADLS account that Prophecy will use by default for writing data. # Google BigQuery Source: https://docs.prophecy.ai/api-reference/connections/properties/bigquery Properties required to create and update BigQuery connections Properties for BigQuery connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties Authentication type for BigQuery. Currently only `private_key` is supported. GCP project ID where your BigQuery datasets are located. Default dataset where Prophecy will write data. Reference to a Prophecy secret containing the service account key JSON. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for service account keys. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. # Databricks Source: https://docs.prophecy.ai/api-reference/connections/properties/databricks Properties required to create and update Databricks connections Properties for Databricks connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties Review the required connection properties for the authentication method you are using. Authentication method for your Databricks connection. Must be `pat` for Personal Access Token authentication. JDBC connection URL for the Databricks SQL warehouse. Format: `jdbc:databricks://{workspace - url}:{port}/{database};transportMode=http;ssl=1;AuthMech=3;httpPath={http - path};` Unity Catalog catalog name for the default write location. Schema name within the catalog for the default write location. Reference to a Prophecy secret containing your Databricks Personal Access Token. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for Personal Access Tokens. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Knowledge graph configuration settings for the connection. See the [Knowledge graph indexer](/api-reference/connections/properties/kgconfig) properties for more information. Authentication method for your Databricks connection. Must be `oauth` for OAuth authentication. JDBC connection URL for the Databricks SQL warehouse. Format: `jdbc:databricks://{workspace - url}:{port}/{database};transportMode=http;ssl=1;AuthMech=3;httpPath={http - path};` Unity Catalog catalog name for the default write location. Schema name within the catalog for the default write location. Provider name. Must be `databricks`. OAuth authentication type. Must be `u2m` for user-to-machine authentication. OAuth app registration ID [configured in Prophecy](/data-analysis/administration/management/cluster-admin-settings/oauth-setup) for Databricks. Knowledge graph configuration settings for the connection. See the [Knowledge graph indexer](/api-reference/connections/properties/kgconfig) properties for more information. Authentication method for your Databricks connection. Must be `oauth` for OAuth authentication. JDBC connection URL for the Databricks SQL warehouse. Format: `jdbc:databricks://{workspace - url}:{port}/{database};transportMode=http;ssl=1;AuthMech=3;httpPath={http - path};` Unity Catalog catalog name for the default write location. Schema name within the catalog for the default write location. Provider name. Must be `databricks`. OAuth authentication type. Must be `m2m` for machine-to-machine authentication. Your OAuth client ID. Reference to a Prophecy secret containing the OAuth client secret. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for client secrets. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Knowledge graph configuration settings for the connection. See the [Knowledge graph indexer](/api-reference/connections/properties/kgconfig) properties for more information. # Google Cloud Storage connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/gcs Properties required to create and update GCS connections Properties for Google Cloud Storage (GCS) connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties Name of the GCS bucket where Prophecy will write data by default. Reference to a Prophecy secret containing the service account key JSON. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for service account keys. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. # SAP HANA Source: https://docs.prophecy.ai/api-reference/connections/properties/hana Properties required to create and update SAP HANA connections Properties for SAP HANA connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. There are two authentication methods supported for SAP HANA connections: * Username/Password authentication. * User Key authentication. Review the required connection properties for the authentication method you are using. ## Properties Use`pwd` for Username/Password authentication. Hostname or IP address of the SAP HANA server. Port number for the SAP HANA server connection. Username for authenticating to the SAP HANA database. Reference to a Prophecy secret containing the password for authentication. Required for `pwd` authentication. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for passwords. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Use `userkey` for User Key authentication. Reference to a Prophecy secret containing the user key for authentication. Required when `authType` is `userkey`. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for user keys. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. # Knowledge graph indexer connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/kgconfig Properties for configuring the knowledge graph indexer for a connection The `kgConfig` object configures [knowledge graph indexer](/data-analysis/ai/knowledge-graph/indexer) settings for connections. Properties vary by connector type and authentication method. ## General properties Whether the knowledge graph is enabled for this connection. Whether to use existing connector authentication for knowledge graph operations. * When `true`, the knowledge graph uses the connection's authentication. * When `false`, the knowledge graph uses separate authentication specified in `kgConfig.authProperties`. The value can only be `false` when the connection's authentication method is **OAuth**. Knowledge graph scheduling configuration. Time zone for knowledge graph scheduling. Format: `America/New_York` or other IANA time zone identifier. Cron expression defining the schedule for knowledge graph operations. Format follows standard cron syntax. Example: `0 0 18 ? * 1,3 *` for 6 PM on Mondays and Wednesdays. Authentication properties for knowledge graph operations. These properties can differ from the main connection's authentication settings. Properties vary by OAuth type (`u2m` or `m2m`). Authentication type for knowledge graph operations. Must be `oauth`. OAuth authentication type for knowledge graph operations. Must be `u2m`. OAuth app registration ID [configured in Prophecy](/data-analysis/administration/management/cluster-admin-settings/oauth-setup). Authentication type for knowledge graph operations. Must be `oauth`. OAuth authentication type for knowledge graph operations. Must be `m2m`. OAuth client ID for knowledge graph operations. Reference to a Prophecy secret containing the OAuth client secret. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for client secrets. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. # MongoDB connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/mongodb Properties required to create and update MongoDB connections Properties for MongoDB connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties MongoDB connection protocol. Use `mongodb` for standard connections or `mongodb+srv` for connections using DNS seed list format (typically used with MongoDB Atlas). MongoDB server hostname or fully qualified domain name. Username for authenticating to the MongoDB server. Reference to a Prophecy secret containing the MongoDB password. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for passwords. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Name of the MongoDB database where Prophecy will write data by default. Name of the collection where Prophecy will write data by default. # Microsoft SQL Server Source: https://docs.prophecy.ai/api-reference/connections/properties/mssql Properties required to create and update MSSQL connections Properties for Microsoft SQL Server (MSSQL) connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties MSSQL server hostname or fully qualified domain name. Port number for the MSSQL server connection. Typically `1433` for SQL Server connections. Username for authenticating to the MSSQL server. Reference to a Prophecy secret containing the MSSQL password. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for passwords. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Knowledge graph configuration settings for the connection. See the [Knowledge graph indexer](/api-reference/connections/properties/kgconfig) properties for more information. # Microsoft OneDrive connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/onedrive Properties required to create and update OneDrive connections Properties for Microsoft OneDrive connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. Consider the following when creating a OneDrive connection: * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * You cannot create a OneDrive connection using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), unless the required secrets are already created in Prophecy. This is because this connection type requires more than one secret, and the create fabric API only supports creating one secret at a time. ## Properties Azure Active Directory tenant ID. Azure Active Directory application (client) ID. Reference to a Prophecy secret containing the Azure client secret. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for client secrets. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Reference to a Prophecy secret containing the OneDrive username. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for usernames. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. # Oracle connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/oracle Properties required to create and update Oracle connections Properties for Oracle connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties Oracle server hostname or fully qualified domain name. Port number for the Oracle server connection. Typically `1521` for Oracle database connections. Username for authenticating to the Oracle server. Reference to a Prophecy secret containing the Oracle password. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for passwords. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Name of the Oracle database to connect to. Knowledge graph configuration settings for the connection. See the [Knowledge graph indexer](/api-reference/connections/properties/kgconfig) properties for more information. # Postgres Source: https://docs.prophecy.ai/api-reference/connections/properties/postgres Properties required to create and update PostgreSQL connections Properties for PostgreSQL connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties PostgreSQL server hostname or fully qualified domain name. Port number for the PostgreSQL server connection. Typically `5432` for PostgreSQL. Name of the PostgreSQL database to connect to. Username for authenticating to the PostgreSQL server. Reference to a Prophecy secret containing the PostgreSQL password. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for passwords. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Knowledge graph configuration settings for the connection. See the [Knowledge graph indexer](/api-reference/connections/properties/kgconfig) properties for more information. # Microsoft Power BI connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/power-bi Properties required to create and update Power BI connections Properties for Microsoft Power BI connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties Azure Active Directory tenant ID. Azure Active Directory application (client) ID. Reference to a Prophecy secret containing the Azure client secret. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for client secrets. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. # Amazon Redshift connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/redshift Properties required to create and update Redshift connections Properties for Amazon Redshift connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties Redshift cluster endpoint hostname or fully qualified domain name. Example: `cluster-name.us-east-1.redshift.amazonaws.com` Port number for the Redshift cluster connection. Typically `5439` for Redshift database connections. Username for authenticating to the Redshift cluster. Reference to a Prophecy secret containing the Redshift password. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy` for Prophecy secrets. Secret sub-type. Must be `text` for passwords. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Use the [list secrets](/api-reference/secrets/list-secrets-per-fabric) endpoint to find the secret ID. The name of the secret stored in Prophecy. Name of the Redshift database to connect to. Knowledge graph configuration settings for the connection. See the [Knowledge graph indexer](/api-reference/connections/properties/kgconfig) properties for more information. # Amazon S3 connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/s3 Properties required to create and update S3 connections Properties for Amazon S3 connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties AWS region where the S3 bucket is located. Example: `us-east-1`, `ap-south-1`, `eu-west-1`. Name of the S3 bucket to connect to. AWS access key ID for authenticating to S3. Reference to a Prophecy secret containing the AWS secret access key. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for secret access keys. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Knowledge graph configuration settings for the connection. See the [Knowledge graph indexer](/api-reference/connections/properties/kgconfig) properties for more information. # Salesforce connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/salesforce Properties required to create and update Salesforce connections Properties for Salesforce connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. Consider the following when creating a Salesforce connection: * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * You cannot create a Salesforce connection using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), unless the required secrets are already created in Prophecy. This is because this connection type requires more than one secret, and the create fabric API only supports creating one secret at a time. ## Properties Salesforce instance URL. For production environments, use the format `https://{instance}.my.salesforce.com` For sandbox environments, use the format `https://{instance}.sandbox.my.salesforce.com` Reference to a Prophecy secret containing the Salesforce username. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for usernames. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Reference to a Prophecy secret containing the Salesforce password. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for passwords. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Reference to a Prophecy secret containing the Salesforce security token. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for security tokens. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Salesforce API version. # SFTP connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/sftp Properties required to create and update SFTP connections Properties for SFTP connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties SFTP server hostname or fully qualified domain name. Port number for the SFTP server connection. Username for authenticating to the SFTP server. Authentication method for SFTP. Supported values: * `password` for password authentication. * `private_key` for private key authentication. Reference to a Prophecy secret containing the SFTP password. Required when `authMethod` is `password`. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for passwords. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Base64-encoded private key for SFTP authentication. Required when `authMethod` is `private_key`. The private key should be in OpenSSH format and encoded as base64. Knowledge graph configuration settings for the connection. See the [Knowledge graph indexer](/api-reference/connections/properties/kgconfig) properties for more information. # Microsoft SharePoint connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/sharepoint Properties required to create and update SharePoint connections Properties for SharePoint connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties SharePoint site URL. Format: `https://{tenant}.sharepoint.com/sites/{sitename}` Azure Active Directory tenant ID. Azure Active Directory application (client) ID. Reference to a Prophecy secret containing the Azure client secret. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for client secrets. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. # Smartsheet connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/smartsheet Properties required to create and update Smartsheet connections Properties for Smartsheet connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties Reference to a Prophecy secret containing the Smartsheet access token. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for access tokens. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. # SMTP connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/smpt Properties required to create and update SMTP connections Properties for SMTP connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties SMTP server hostname or fully qualified domain name. Example: `smtp.gmail.com` Port number for the SMTP server connection. for SSL. Username for authenticating to the SMTP server. Reference to a Prophecy secret containing the SMTP password. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for passwords. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. # Snowflake connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/snowflake Properties required to create and update Snowflake connections Properties for Snowflake connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties Authentication type for Snowflake. Use `pwd` for password authentication. Snowflake account identifier. Format: `{account_identifier}.{region}.{cloud_provider}.snowflakecomputing.com` Snowflake username for authentication. Reference to a Prophecy secret containing the Snowflake password. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for passwords. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Snowflake warehouse where Prophecy will write data by default. Database where Prophecy will write data by default. Schema where Prophecy will write data by default. Snowflake role to use for the connection. Determines the permissions and access level for operations. # Azure Synapse connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/synapse Properties required to create and update Synapse connections Properties for Synapse connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties Azure Synapse Analytics server hostname or fully qualified domain name. Port number for the Synapse server connection. Typically `1433` for SQL Server connections. Username for authenticating to the Synapse server. Reference to a Prophecy secret containing the password. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for passwords. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. Name of the database on the Synapse server that Prophecy will use by default for writing data. # Tableau connection properties API Source: https://docs.prophecy.ai/api-reference/connections/properties/tableau Properties required to create and update Tableau connections Properties for Tableau connections are specified in the `properties` object when creating or updating a connection. All properties are nested within the connection request body. ## Properties Tableau server URL. Name of the Tableau site to connect to. Name of the personal access token that you created in Tableau. Reference to a Prophecy secret containing the Tableau personal access token value. * When using the [create connection API](/api-reference/connections/add-connection-to-fabric), the secret must be stored in Prophecy before creating the connection. * When using the [create fabric API](/api-reference/fabrics/create-a-new-fabric), replace this whole object with the string `{{SECRET}}`, which references the secret being created in parallel to the connection. Secret kind. Must be `prophecy`. Secret sub-type. Must be `text` for token values. Type of reference. Must be `secret`. Properties identifying the secret by ID and name. The ID of the secret stored in Prophecy. Find the ID using the [List secrets per fabric API](/api-reference/secrets/list-secrets-per-fabric). The name of the secret stored in Prophecy. # Retrieve connection details Source: https://docs.prophecy.ai/api-reference/connections/retrieve-connection-details get /api/orchestration/fabric/{fabricId}/connection/name/{connectionName} Retrieve details for a specific connection by its name # Update connection Source: https://docs.prophecy.ai/api-reference/connections/update-connection put /api/orchestration/fabric/{fabricId}/connection/name/{connectionName} Update the properties of a specific connection # Create a new fabric Source: https://docs.prophecy.ai/api-reference/fabrics/create-a-new-fabric post /api/orchestration/fabric Create a Prophecy fabric with optional secret and connection You can only add one secret and one connection in the same request. In most cases, the connection will require a secret, so these will be created together. The `connection.properties` object and `secret.properties` object depend on the connection type and secret type. Find specific properties in the following API documentation sections: * **Connections > Properties** * **Secrets > Properties** # Delete a fabric Source: https://docs.prophecy.ai/api-reference/fabrics/delete-a-fabric delete /api/orchestration/fabric/{fabricId} Delete a fabric by its ID # Get fabric details Source: https://docs.prophecy.ai/api-reference/fabrics/get-fabric-details get /api/orchestration/fabric/{fabricId} Retrieve details for a specific fabric by its ID # Update an existing fabric Source: https://docs.prophecy.ai/api-reference/fabrics/update-an-existing-fabric put /api/orchestration/fabric/{fabricId} Update the name or description of an existing fabric Only the name and description of the fabric can be updated. # Prophecy API Source: https://docs.prophecy.ai/api-reference/introduction Use APIs to interact with your Prophecy deployment The Prophecy API provides programmatic access to your Prophecy deployment through REST endpoints. These endpoints accept JSON payloads and return structured responses, enabling integration with CI/CD pipelines, monitoring systems, and custom automation scripts. All requests require authentication via a Personal Access Token in the request headers. Prophecy has an additional set of APIs that you can request through Prophecy Support. Reach out to Support to learn more. ## Endpoints Each endpoint uses your Prophecy environment URL as the base URL. Replace the base URL with your environment URL for Dedicated SaaS deployments. ## Access tokens Prophecy uses your Personal Access Token (PAT) for authentication with our API servers. To manage your access tokens: 1. Open your Prophecy environment. 2. At the bottom of the left sidebar, click the **...** menu. 3. Click the gear icon to open Settings. 4. Navigate to the **Access Tokens** tab. Here, you will see any PATs you have previously created. You can also find information such as creation date, expiration status, and time of last use. Access tokens are per user. Therefore, API permissions are scoped to your user permissions. ### Generate an access token To generate a new access token: 1. Open the **Access Tokens** tab in Settings. 2. Click **Generate Token**. 3. Name your token. 4. Choose an expiration date from the dropdown menu. 5. Click **Create**. 6. Copy your newly-generated token before closing the dialog. You will not be able to access the token value again. # Get pipeline run status Source: https://docs.prophecy.ai/api-reference/pipelines/get-pipeline-run-status get /api/trigger/pipeline/{runId} Get the status of a triggered pipeline run, including an error message if the pipeline run fails. # Run data tests Source: https://docs.prophecy.ai/api-reference/pipelines/run-data-tests post /api/orchestration/tests/run Execute existing data tests for a project Use this endpoint to run [data tests](/data-analysis/development/tests/test-comparison). Column tests, table tests, and project tests can be run together in a single API call. Model tests must be run in a separate API call and cannot be mixed with other test types. ## Requirements To run data tests for a project using the API, you need to: * Publish the project to a [Prophecy fabric](/data-analysis/environment/fabrics/prophecy-fabrics). * Set up data tests in the Prophecy [Studio](/data-analysis/development/studio/studio). This includes both creating tests and assigning them to tables or models. * Ensure that any tables you want to test have data. You'll have to run your pipelines once to populate target tables. # Trigger pipeline run Source: https://docs.prophecy.ai/api-reference/pipelines/trigger-pipeline-run post /api/trigger/pipeline Trigger the execution of a pipeline in a deployed project. This is useful when you want to automate pipeline runs without relying on Prophecy's built-in scheduling. ## Requirements To run the API, the target pipeline must: * Run on a [Prophecy fabric](/data-analysis/environment/fabrics/prophecy-fabrics). * Be part of a published project. In other words, the pipeline must be deployed. # Deploy project Source: https://docs.prophecy.ai/api-reference/projects/deploy-project post /api/deploy/project Deploy projects with custom pipeline and project configurations to specific fabrics. Use this API to automate project deployment with external CI/CD tools or deploy the same project to different environments with different configuration values. ## Requirements To run the API, you need to: * Reference a fabric that exists in your environment. * Use an existing Git tag that defines the project version. Unlike in the Prophecy UI where Git tags are created automatically during publish, when using this API you must create the Git tag externally. The tag must follow the exact format `{projectName}/{version}`. ## Parameter Precedence Configuration parameters follow this precedence order: 1. Pipeline parameter overrides (highest priority) 2. Project configuration overrides 3. Default pipeline parameters 4. Default project configuration (lowest priority) # Add secret to fabric Source: https://docs.prophecy.ai/api-reference/secrets/add-secret-to-fabric post /api/orchestration/fabric/{fabricId}/secret Create a new secret in a specific fabric # Delete secret Source: https://docs.prophecy.ai/api-reference/secrets/delete-secret delete /api/orchestration/fabric/{fabricId}/secret/id/{secretId} Delete a specific secret by its ID # List secrets per fabric Source: https://docs.prophecy.ai/api-reference/secrets/list-secrets-per-fabric get /api/orchestration/fabric/{fabricId}/secret List all secrets and their details from a specific fabric # Binary Source: https://docs.prophecy.ai/api-reference/secrets/properties/binary API properties for binary secrets Properties for binary secrets are specified in the `properties` object when creating or updating a secret. Binary secrets store base64-encoded binary data such as files, certificates, or other binary content. ## Properties The key name for the secret. This name is used to reference the secret in connections and other resources. The secret value as a base64-encoded string. The binary data must be encoded to base64 before being sent in the request. # M2M OAuth Source: https://docs.prophecy.ai/api-reference/secrets/properties/oauth API properties for OAuth secrets Properties for M2M OAuth secrets are specified in the `properties` object when creating or updating a secret. OAuth secrets store OAuth 2.0 client credentials including client ID, client secret, and authorization URL. ## Properties The key name for the secret. This name is used to reference the secret in connections and other resources. OAuth 2.0 credentials object containing client authentication details. The OAuth 2.0 client identifier. This is the public identifier for your OAuth application. The OAuth 2.0 client secret. This is the private credential used to authenticate your application. The authorization URL for the OAuth provider. This is the base URL where OAuth authentication requests are sent. OAuth 2.0 scope string specifying the permissions requested from the OAuth provider. Multiple scopes should be space-separated. # Text Source: https://docs.prophecy.ai/api-reference/secrets/properties/text API properties for text secrets Properties for text secrets are specified in the `properties` object when creating or updating a secret. Text secrets store plain text values such as API keys, service account JSON, or other string-based credentials. ## Properties The key name for the secret. This name is used to reference the secret in connections and other resources. The secret value as a plain text string. For service account keys, this should be the complete JSON string. # Username Password Source: https://docs.prophecy.ai/api-reference/secrets/properties/username-password API properties for username-password secrets Properties for username-password secrets are specified in the `properties` object when creating or updating a secret. Username-password secrets store basic authentication credentials with a username and password pair. ## Properties The key name for the secret. This name is used to reference the secret in connections and other resources. Username and password credentials object. The username for basic authentication. The password for basic authentication. # Retrieve secret details Source: https://docs.prophecy.ai/api-reference/secrets/retrieve-secret-details get /api/orchestration/fabric/{fabricId}/secret/id/{secretId} # Update secret Source: https://docs.prophecy.ai/api-reference/secrets/update-secret put /api/orchestration/fabric/{fabricId}/secret/id/{secretId} Update the properties of a specific secret Open the secret in the Prophecy UI to see which properties can be updated. # Prophecy Agent overview Source: https://docs.prophecy.ai/data-analysis/ai/agent/agent A development agent for building and modifying data pipelines using natural language The Prophecy Agent uses natural language to build, modify, and interpret pipelines, analyses, documentation, and related project artifacts. When you describe your intent, the Agent inspects your active project, searches metadata, retrieves sample data where permitted, and generates or updates transformations directly within the project. Importing from Alteryx or other platforms? You can start by [importing these intro Prophecy](https://docs.prophecy.ai/import-tool). Then use the Agent to modify and extend these workflows. ## Why use the Agent? The Agent accelerates pipeline development by reducing the manual steps required to build and modify data workflows. Instead of navigating multiple interfaces, you describe intent. The Agent translates that intent into concrete, inspectable project changes within your active environment. The Agent assists development; it does not autonomously deploy or promote changes. ## Scope and control The Agent operates within the currently active project and under the permissions of the user who invoked it. * All SQL execution and pipeline runs occur under your existing credentials. * All changes remain visible, versioned, and editable within Prophecy. Prophecy includes specialized agents for transformation, harmonization, and documentation workflows. The Transform Agent runs separately from the Documentation and Harmonization Agents. To use the Harmonization Agent or Documentation Agent, you must disable the Transform Agent. (Harmonization and Documentation can be used without disabling one another.) * [Transform Agent](/data-analysis/ai/agent/transform) (default) β€” Builds and modifies pipelines, analyses, and documentation using natural language. * [Harmonization Agent](/data-analysis/ai/agent/harmonization/overview) β€” Automates mapping source data to a defined Common Data Model (CDM). * [Documentation Agent](/data-analysis/ai/agent/documentation/documentation) β€” Generates complete project and pipeline documentation. ## Workflow Agent workflows vary depending on the active agent. ### Transform Agent workflow (default) The Transform Agent focuses on building and modifying [pipelines](/data-analysis/), [analyses](/data-analysis/ai/agent/analyses), and related project artifacts within the active project. When you submit a request, the Agent: * Interprets your intent. * Inspects the current project graph and schema metadata. * Generates or modifies pipeline logic. * Validates transformations and project structure. * Surfaces changes for review. The Transform Agent supports two primary intents throughout the pipeline development lifecycle: * [Exploration](/data-analysis/ai/agent/explore) for discovering and understanding data sources. * [Transformation](/data-analysis/ai/agent/transform) for building and modifying pipeline logic. Use the Agent to search your data warehouse, preview datasets, and validate data quality before building transformations. During transformation, describe data operations in natural language to generate gems and modify pipeline logic. ### Harmonization workflow The [Harmonization Agent](/data-analysis/ai/agent/harmonization/overview) focuses on mapping source schemas to a defined Common Data Model (CDM). The workflow typically includes: * Defining or selecting a CDM. * Generating source-to-target mappings. * Reviewing confidence indicators and data quality tests. ## Agent architecture The Prophecy Agent operates within Prophecy’s structured execution environment and is equipped with tools to query warehouses, compile code, run pipelines, and modify project artifacts directly. Rather than generating text alone, the Agent uses controlled internal tools to inspect, modify, and validate project assets directly within the active project. This enables the Agent to reason over: * Your project's structure (pipelines, analyses, datasets, documentation). * Schema metadata and dataset relationships. * Version-controlled project artifacts. ### Tool-based execution The Agent uses a set of internal tools to take concrete actions inside your project, including: * Querying your connected warehouse. * Retrieving metadata and sample data. * Compiling and validating code. * Running pipelines. * Creating and editing project artifacts. When the Agent executes SQL or runs a pipeline, it does so under the permissions of the user who invoked it. All warehouse interactions respect existing access controls and governance policies. Before applying changes, the Agent compiles and validates generated logic against your project structure to detect syntax errors, broken references, or incompatible transformations. ### Schema and lineage awareness The Agent understands your project’s schema metadata and dataset relationships. When generating or modifying transformations, it: * Inspects upstream schema definitions. * Validates generated changes against your project structure before applying them. * Surfaces potential conflicts or mismatches when detected. The Agent assists with structural consistency but does not override warehouse-level schema constraints. ### Project boundary All Agent activity is strictly contained within the active project and cannot affect external projects or platform-level configuration. It can: * Create and refactor pipelines. * Generate and update analyses. * Modify documentation. * Edit datasets and other project artifacts. ### Single-agent editing Only one Agent session can edit a project at a time. Collaborative agent editing across multiple users is not supported. ### Human in the loop The Agent generates and can apply changes inside your project, but you remain in control. All modifications are visible in the editor and can be inspected, edited, or reverted. You can: * Inspect every transformation. * Validate data samples. * Modify generated logic. * Re-run pipelines as needed. The Agent does not replace review. Users remain responsible for validating logic, data correctness, and production readiness. ```mermaid theme={null} graph TD A[Start Project] --> B[Find Data Sources] B -->|Exploration Agent| C[Preview & Validate Data] C -->|Exploration Agent| D[Add Sources to Pipeline] D --> E[Build Pipeline Logic] E -->|Transformation Agent| F[Review & Refine] F -->|Human in the loop| G{Complete?} G -->|No| E G -->|Yes| H[Generate Documentation] H -->|Documentation Agent| I[Deploy & Maintain] ``` ## Available features Use table metadataβ€”table names, schemas, owners, and tagsβ€”to locate datasets from your fabric. When you don't know the exact table name or work with many tables, searching through metadata eliminates guesswork and reduces time spent browsing schema lists. You can find data from your connected data warehouse, cloud storage, reporting platforms, and other sources. Preview data samples and visualizations before selecting datasets for your pipeline. Understanding column structures, data patterns, and potential quality issues helps you make informed decisions about which tables to use. Generate complete pipelines from a single description when requirements are clear, or build incrementally gem by gem, adding one transformation at a time in a linear sequence. Instead of manually dragging gems onto the canvas and configuring each step, describe your goal and let the Agent handle the setup. The Harmonization Agent lets you define a Common Data Model (CDM) and generate mappings that conform input data to the CDM. Learn more in the [Harmonization](/data-analysis/ai/agent/harmonization/overview) documentation. Clean up existing pipelines by removing unnecessary transformations and consolidating logic. Easier-to-maintain pipelines execute faster, and clearer logic helps teammates understand your work without deciphering complex transformation chains. Get a high-level explanation of what a pipeline does and which datasets it uses without reading through every transformation. Understanding existing pipelines quickly helps with onboarding, while documenting your own work makes it easier for others to use and modify later. ## Features in progress Features in progress are subject to the same project-scope and permission constraints as existing capabilities. Provide sample data to the Agent and retrieve similar datasets from your fabric. Iterate on granular aspects of the pipeline while ensuring previous work remains intact. Identify opportunities to improve pipeline performance using methods such as query optimizations, caching, or simplified joins. Ask the Agent to start crawling sources to ensure the most up-to-date metadata in the knowledge graph. Generate tests to validate data quality and catch errors before pipelines run in production. Find relevant packages that help build out your pipeline for your specific use case. Automate moving pipelines from development to production. Track changes, create branches, and manage pipeline versions through the Agent interface. # Agent governance and data access Source: https://docs.prophecy.ai/data-analysis/ai/agent/agent-governance How Prophecy agents interact with your data and respects access controls Prophecy agents operate within your selected deployment environment and respects existing access controls. The agents are assistive rather than autonomous: they do not have autonomy to connect to external systems or perform any action without approval. Prophecy agents are considered limited risk according to the EU AI Act. This page explains how data is accessed, executed, and stored when using agents. ## Key principles * Agents generate transformation logic and suggestions, but you review and approve all changes before pipelines are run in production. * Agents operate only within your Prophecy deployment and connected data platforms. * You can inspect all generated transformations in visual, SQL, or documentation form before execution. * You maintain full control over whether changes are applied. ## Storage and environment Agents run inside your Prophecy deployment environment. Prophecy stores: * Agent chat history. * Project metadata (pipelines, analyses, schemas, configurations). * Version-controlled project artifacts. These items remain within your Prophecy deployment environment. ## LLM interaction When generating suggestions, agents may include project metadata in prompts sent to the configured language model. This metadata may include: * Table and column names * Schemas and data types * Transformation code within the current project The agent does not include sample data by default. Using sample data can only be enabled by an administrator. ## Execution model Agents generate transformation logic within your Prophecy deployment environment. When validating transformations or running pipelines, execution occurs deterministically on your connected data platform and under the permissions of the user who invoked the agent. All reads and writes respect existing warehouse access controls, role-based permissions, and organizational governance policies. Agents do not independently access external systems or perform actions outside your configured data platforms. # Analyses Source: https://docs.prophecy.ai/data-analysis/ai/agent/analyses Structured analytical views generated from pipeline outputs Analyses are structured analytical views associated with a pipeline. They appear in the Project Browser beneath their parent pipeline and are represented with a bar graph icon. An Analysis presents business insights derived from pipeline outputs, including visualizations, recommendations, and supporting data details. ## Relationship to pipelines Analyses are generated within the context of a pipeline and depend on that pipeline's outputs. Pipelines prepare and transform data. Analyses interpret and present insights from those outputs for decision-making. Because Analyses are tied to pipelines: * They reflect the logic defined in the underlying transformations. * Updates to pipeline logic affect Analysis results. * A single pipeline can support multiple Analyses focused on different business questions. ## What an Analysis includes An Analysis may include: * Key insights summarizing findings. * Bar charts and other visualizations. * Key recommendations based on results. * Supporting metrics and data details. * Interactive configuration controls. These elements are structured and versioned as part of the project. Example of analysis ## Creating Analyses with the Agent Analyses are generated by the Agent in response to business questions. When you ask a question, the Agent can: * Identify relevant datasets and pipeline outputs. * Generate or modify SQL models. * Create an Analysis beneath the pipeline. * Populate the Analysis with insights, charts, and recommendations. In some cases, the Agent may suggest creating an Analysis dashboard for the pipeline. All generated assets are saved within the project. You can inspect and refine the SQL logic, adjust visualizations, and modify recommendations as needed. ## Reuse and refinement Once created, an Analysis can be reused and refined without rebuilding the underlying pipeline. You can modify an Analysis by: * Updating filters or scoring thresholds in the parent pipeline. * Adjusting business rules. * Editing insight summaries. * Adding or removing visualizations. * Rearranging the order of sections within the dashboard. Changes to the Analysis do not require recreating the pipeline. Because the logic and configuration are stored within the project, Analyses can evolve over time while remaining versioned and editable. This allows you to iterate on insights, test different segments or scenarios, and adapt recommendations as business needs change. # Export chat session info Source: https://docs.prophecy.ai/data-analysis/ai/agent/chat/export Export session information from Agent chat for troubleshooting and support Export session information from your Agent chat to share with support or engineering teams when troubleshooting issues. The exported data includes session identifiers, environment details, and version information that helps reproduce your chat state. ## Steps 1. Open chat history. 2. Click **Copy Session Details**. 3. The session information is copied to your clipboard as JSON. ## Session info structure The exported session info contains the following fields: Unique identifier for the current chat session. Identifier for the specific chat conversation. The Prophecy application hostname where the session is active. The commit hash or version identifier of the Prophecy build running in your environment. Unix timestamp in milliseconds indicating when the session info was exported. The Prophecy version number running in your environment. The identifier for the project for this chat. The email for the chat user. ## Example output When you export session info, the clipboard contains JSON in this format: ```json theme={null} { "sessionId": "861716af-6414-407d-97d1-d7afa1866558", "conversationId": "x2G3r2uxsPX8CpzRDo_Us", "hostname": "app.prophecy.io", "commit": "4.2.1.4", "timestamp": 1763401148947, "version": "4.2.1.4", "projectId": "31709", "userEmail": "user.drew@prophecy.io" } ``` Paste the exported session info into a support ticket or share it with your engineering team to help diagnose issues specific to your chat session and environment. # @ mentions Source: https://docs.prophecy.ai/data-analysis/ai/agent/chat/mentions Use @ mentions to refer to specific entities in your project Reference specific entities in your project using `@` mentions. When you type `@` or click the `@` button in the chat footer, Prophecy displays an autocomplete dropdown with matching entities from project. Use `@` mentions to eliminate ambiguity when your prompt could refer to multiple entities. The Agent uses the exact entity you mention rather than inferring from context, which reduces errors in complex pipelines. ## How to use @ mentions Type `@` in the chat input field, or click the `@` button in the chat footer. An autocomplete dropdown appears with suggested entities. Type additional characters to filter the suggestions, or scroll through the list. Each entity type displays with a distinct icon to help you identify it quickly. Click an entity from the dropdown, or use arrow keys to navigate and press Enter. The `@` mention is inserted into your prompt. ## Supported entities The autocomplete dropdown includes entities from your current project. * Tables: Reference tables in your SQL warehouse using full table names, including database and schema information (for example, `database.schema.table`). This is useful when you want to add a table to your pipeline or ask the Agent to provide sample data from a table. Only tables indexed by the knowledge graph appear in the autocomplete suggestions. If a table doesn't appear, ensure it's been indexed through the [knowledge graph indexer](/data-analysis/ai/knowledge-graph/indexer). * Gems: Reference specific gems in your current pipeline by their label. Use the [gem label](/data-analysis/gems/gems#gem-instance) that identifies the instance, not the gem type. This is useful when you want to add a transformation to your pipeline or ask the Agent to provide sample data from a gem. * Documentation templates: Reference documentation templates in your project when generating documentation. Templates must exist in your project to appear in suggestions. Use with prompts like `Generate pipeline documentation using @custom_template`. ## Examples ### Reference tables ``` How many records are in the @transactions table? ``` ``` Show me the schema for @energy_production_metrics ``` ### Reference gems ``` Aggregate the @orders_cleaned gem to sum order quantity per month ``` ``` Add filter after @last_aggregation_gem to show only records from 2024 ``` ### Reference documentation templates ``` Generate pipeline documentation using @custom_template ``` # Upload files Source: https://docs.prophecy.ai/data-analysis/ai/agent/chat/upload-files Upload files for data pipelines, analysis, or chat context When you upload a file to the Agent chat, you can choose one of two options: | Option | Use when | | ----------------------- | ------------------------------------------------------------------------------------------------------------------- | | **Process as Table** | Convert the file into a table in your SQL warehouse so the Agent can use it in data pipelines and transformations. | | **Attach as Reference** | Keep the file in its original format so the Agent can use it as chat context, documentation, or reference material. | The remainder of this page describes **Process as Table** uploads. When you choose **Process as Table**, Prophecy uploads the file to the SQL warehouse configured in your fabric and converts it into a structured table that the Agent can read and use in transformations. Each file can be up to 30 MB, and the total size of all files in a single upload can be up to 100 MB. Uploaded files are scanned for safety before they are processed. Prophecy rejects uploads if they contain disallowed file types, executable content, or archives that exceed size, nesting, or entry-count limits. Rejected uploads show an error message describing the reason. You can upload the following file formats: * CSV * Excel (XLS, XLSX) * JSON * Parquet * XML ## Upload a file Click the **paperclip** icon in the chat footer or drag a file onto the pipeline canvas. When prompted, choose **Process as Table** to continue with the workflow described on this page. Choose the file you want to upload from your local filesystem and drag it into chat. The upload dialog opens automatically. Determine where the file should be stored in the SQL warehouse: * Choose the database. * Choose the schema. * Choose the table name. You can select an existing table or create a new one. Click **Next** to continue. If you select an existing table, Prophecy deletes and recreates the table with your uploaded file data. Review the inferred schema (column names and data types): * Keep the auto-detected schema, or modify column names and data types as needed. * Select any table properties you want to configure. * Click **Next** to proceed. Excel files (XLS, XLSX) expose additional options at this step: | Field | Purpose | | -------------------------- | ------------------------------------------------------------------------------------------------ | | **Read mode** | How the workbook is read: a single sheet, or a union of multiple sheets combined into one table. | | **Sheet Name** | Which sheet to read from. When reading a union of multiple sheets, select each sheet to include. | | **Enter Cell Range** | Restrict extraction to a specific range, for example `A1:`. | | **Header** | Treat the first row of the range as column headers. | | **Allow Undefined Rows** | Include rows that fall outside the detected schema. | | **Allow Incomplete Rows** | Include rows missing values in some columns. | | **Enable Schema Merging** | Merge schemas across sheets or tables where they differ. | | **Ignore Cell Formatting** | Read cell values without applying Excel's display formatting. | After you configure these options, click **Infer Schema** to populate the column list. Column names, types (for example, String or Bigint), and other metadata are editable. You can also mark multiple regions of an Excel workbook as separate tables. Each marked table becomes its own output, and you can edit its schema, set drift-handling behavior, or unmark it independently of the others. Review the table preview to verify the data structure and content. If the preview looks correct, click **Done** to complete the upload. For Excel files, the preview grid loads additional rows and columns as you scroll, rather than fetching the entire sheet at once. Merged cells and blank rows in the source sheet are handled automatically. The file is converted to a table in your SQL warehouse. ## Make the table available to the Agent For the Agent to discover the table, you need to add it to the knowledge graph. After uploading a file, the [knowledge graph indexer](/data-analysis/ai/knowledge-graph/indexer) will automatically index the table. However, you can also manually reindex the knowledge graph, in case something goes wrong. 1. Open the **Environment** tab in the left sidebar. 2. Below your connections, locate the **Missing Tables?** callout. 3. Click **Refresh** to trigger the knowledge graph indexer. Once indexed, the Agent can reference the table using `@` mentions or discover it through natural language queries. The Agent will use the [Table gem](/data-analysis/gems/source-target/table) to read the file into your pipeline. # Use AI chat in projects Source: https://docs.prophecy.ai/data-analysis/ai/agent/chat/using-chat Learn how to work with AI chat in Prophecy projects, including managing conversations, referencing entities, reviewing AI-generated changes, and providing context to the Agent. Prophecy's Agent lets you build and modify pipelines, analyses, and other project entities using natural language. ## What you can do with the Agent You can use the Agent * create or update entities. * ask questions about your project. * reference existing project entities. * upload files for context. * review AI-generated changes. * manage multiple chat sessions ## Send prompts Use the message box at the bottom of the chat panel to send prompts to the Agent. You can ask the Agent to: * create new transformations or analyses * modify existing entities * explain project logic * troubleshoot pipelines * answer questions about project assets Example prompts: * `Create a new transformation that filters failed orders` * `Add a join between customer and transaction tables` * `Explain how this model calculates churn risk` ## Provide context to the Agent You can provide additional context to help the Agent understand your project and requests. ### Upload files Use the attachment icon in the message box to upload files. Uploaded files can help the Agent: * understand schemas * review requirements * reference specifications * analyze supporting documents For more information, see Upload files. ### Reference project entities Use `@` mentions to reference project entities directly in chat. You can reference: * pipelines * models * analyses * project assets Referencing entities gives the Agent additional project context and helps it understand which assets you want to modify or discuss. For more information, see Reference project entities. ## Manage conversations AI chat supports multiple conversations within the same project. ### View chat history Use the **Show history** button at the top of the chat window to view previous conversations. The currently active chat is labeled **Current**. From chat history, you can: * reopen previous chats * rename chats * delete chats * copy session details ### Start a new chat Use **New chat** to begin a separate conversation. Starting a new chat does not remove existing project entities. When starting a new chat, the Agent may reference existing project entities and suggest: * updating existing transformations * creating new analyses * continuing work within the current project ### Rename chats You can rename chats from: * the chat header * the chat history panel To rename a chat: 1. Click the chat name. 2. Enter a new name. 3. Save the updated name. Use descriptive names to organize project conversations. ### Delete chats Delete chats from the chat history panel. Deleting a chat removes that conversation from your project history. Deleting the current chat automatically opens a new chat session. ### Copy session details Use **Copy session details** from chat history to copy information about a conversation. This can help when: * sharing troubleshooting context * collaborating with teammates * reporting issues * documenting AI-generated changes For more information, see Copy chat session details. ## Review AI-generated changes The Agent can modify project entities directly from chat interactions. ### Review modified entities After the Agent makes changes, the chat displays a review area showing which entities were modified. Review changes carefully before continuing work. Modified entities may include: * pipelines * models * analyses * transformations ### Undo the latest AI action The most recent AI-generated change includes an **Undo this** action below the response. Use this action to revert the latest AI-generated modification. The undo option is only available for the most recent response containing changes. Older responses display standard copy actions instead. ## Interactive clarification questions in SQL Copilot The Agent may pause during generation to ask for additional input before continuing. When the Agent needs clarification, it displays a question directly in the chat along with available answer options. After you select a response, the agent continues generation using your choice. This improvement allows the Agent to gather missing information and resolve ambiguities instead of making assumptions or failing when additional context is required. You can also dismiss a question if you do not want to provide an answer. The agent handles skipped questions gracefully and continues with the best available context. ### Confirmation prompts for Agent pipeline execution The Agent requests confirmation before executing pipelines that can write data to your warehouse. When the Agent attempts to run a pipeline, you'll see a prompt identifying the pipeline and warning that the operation may write or overwrite data. You can choose to: * **Allow Once** to approve the current execution * **Always Allow** to automatically approve future executions * **Deny** to prevent the execution This additional confirmation step helps prevent unintended data modifications while giving you control over Agent-initiated pipeline runs. ## Continue work across chats Chats are project-aware and can reference existing entities in your project. Even when starting a new chat, the Agent can: * detect existing project assets * help continue existing workflows * create new entities related to previous work This allows you to organize conversations by task while continuing work within the same project. ## Tips for working with AI chat * Use descriptive chat names for larger projects. * Reference entities with `@` mentions when modifying existing assets. * Upload supporting files to provide additional context. * Review AI-generated changes before continuing work. * Use separate chats for unrelated workflows or experiments. # Copilot Template Language (CTL) Source: https://docs.prophecy.ai/data-analysis/ai/agent/documentation/ctl-reference Find definitions for markers in Copilot Template Language Copilot Template Language (CTL) extends Markdown with special markers that create components in your documentation that standard Markdown doesn't support. These markers enable features like interactive question forms, visual pipeline diagrams, and dynamic content that automatically updates based on your pipeline structure. ## Marker syntax Use CTL markers directly inside your documentation template Markdown file. Markers use this structure: ``` [MARKER_TYPE]() ``` The `MARKER_TYPE` tells the Agent what type of special component to create. Each marker generates a specific type of componentβ€”like interactive forms, visual diagrams, or structured data tablesβ€”that standard Markdown cannot create on its own. For example, when the Agent comes across the `[TRANSFORMATIONS]()` marker in the template file, it will replace the marker with a formatted section listing all transformation steps in your pipeline in the generated documentation file. Each marker may accept zero, one, or multiple arguments. Review the specific marker to see its correct format. ## Marker-template compatibility This table describes which markers are compatible with which template types. | Marker | Pipeline templates | Project templates | | -------------------------------------- | ------------------ | ----------------- | | [OVERVIEW](#overview) | βœ” | βœ” | | [SOURCES](#sources) | βœ” | | | [TARGETS](#targets) | βœ” | | | [TRANSFORMATIONS](#transformations) | βœ” | | | [COPILOT\_QUESTION](#copilot-question) | βœ” | βœ” | | [QUESTION](#question) | βœ” | βœ” | | [PIPELINES](#pipelines) | | βœ” | ## Marker reference The Copilot Template Language includes the following markers. ### Overview ``` [OVERVIEW]() ``` The Agent-generated documentation will have a summary section of the pipeline or project. When the Agent generates the documentation for one pipeline specifically, it will embed an interactive visual diagram of the pipeline directly in the document. ### Sources ``` [SOURCES]() ``` The Agent-generated documentation will have a formatted section that lists all data sources used in the pipeline, including descriptions and column information. This marker does not accept any parameters. This marker is compatible with pipeline-level templates only. ### Targets ``` [TARGETS]() ``` The Agent-generated documentation will have a formatted section that lists all output tables or files created by the pipeline, including descriptions and column information. This marker does not accept any parameters. This marker is compatible with pipeline-level templates only. ### Transformations ``` [TRANSFORMATIONS]() ``` The Agent-generated documentation will have a formatted section with a summary and detailed explanation of each data transformation step in the pipeline. This marker does not accept any parameters. This marker is compatible with pipeline-level templates only. ### Copilot question ``` [COPILOT_QUESTION]('Your question here') ``` The Agent-generated documentation will have an interactive question form in your documentation. The Agent tries to answer the question automatically based on the pipeline or project. If the Agent can't answer, the question appears as a form field for users to fill in after the documentation is generated. This marker accepts one string parameter that contains your question or prompt. For example: `'[COPILOT_QUESTION]('Specify the number of transformation steps in the pipeline.')` The Agent will be able to answer most questions about the project and pipelines. It will not be able to answer questions that require additional context external to the project. ### Question ``` [QUESTION]('Your question here') ``` The Agent-generated documentation will have an interactive question form in your documentation that users must answer manually. Unlike `COPILOT_QUESTION`, the Agent does not try to answer this question. It always appears as a form field for users to complete. This marker accepts one string parameter that contains your question or prompt. For example: `'[QUESTION]('Define the expected outcome of this pipeline')` ### Pipelines ``` [PIPELINES](template=subtemplate.md) ``` The Agent will use a pipeline template to generate documentation and embed that documentation in the higher-level project documentation. The Agent processes the template separately for each pipeline, filling in the details specific to that pipeline. The template must already exist in the project for the Agent to locate it. To learn how to upload a new template, see [Document pipelines](/data-analysis/ai/agent/documentation/documentation). The Agent can only generate one sub-document per marker. Use multiple pipeline markers to incorporate multiple pipeline documentation embeddings. # Generate documentation Source: https://docs.prophecy.ai/data-analysis/ai/agent/documentation/documentation Automatically create comprehensive documentation for project pipelines Use the Prophecy Agent to generate documentation for pipelines or entire projects. The Agent uses [templates](#documentation-templates) to structure the output as Markdown files that describe your data pipelines. You need to upload separate templates for pipeline and project documentation tasks. You can generate documentation for a single pipeline or for an entire project. ## Generate pipeline documentation To generate documentation for a pipeline, follow these steps: 1. Open your pipeline in Prophecy. 2. Click **Chat** in the left sidebar. 3. Prompt the Agent to document your pipeline. For example: `Generate pipeline documentation using @custom_template`. This generates a document for the pipeline that you have opened in your project. If you don't specify a template using an @ mention, the Agent looks for a default template file in your project, which should have the name `pipeline-template.md`. If the default template exists, the Agent uses it. Otherwise, the Agent asks you to upload or specify a template. Ask Agent to generate documentation ## Generate project documentation To generate documentation for a project, including all of its pipelines, follow these steps: 1. Open your project in Prophecy. 2. Open any pipeline or document entity. 3. Click **Chat** in the left sidebar. 4. Prompt the Agent to document your project. For example: `Generate project documentation using @custom_template`. This generates a document that describes your project. If you don't specify a template using an @ mention, the Agent looks for a default template file in your project, which should have the name `project-template.md`. If the default template exists, the Agent uses it. Otherwise, the Agent asks you to upload or specify a template. ## How the Agent generates documentation During generation, the Agent: 1. Displays a preparation checklist with progress. 2. Analyzes pipeline or project components. 3. References the template that defines the document structure. 4. Generates the documentation. 5. Saves the Markdown file to the `documents` directory in your project. View the generated document in **Documents** in the project sidebar. Generated documentation ## Update generated documentation After the Agent generates your document, you can: * Edit the document in the visual editor (recommended). The visual editor displays the rendered Markdown file, including special components like pipeline frames. * Switch to the **Code** tab to edit the Markdown code. In the Code tab, you may see encoded or complex strings that contain embedded data. Don't edit them, as changes will break the visual output. Since the document is stored in your project repository, all changes are version-controlled with your code. ## Documentation templates Templates define how the Agent structures documentation output. Templates use Markdown for standard formatting and [Copilot Template Language (CTL)](/data-analysis/ai/agent/documentation/ctl-reference) markers to create special components that standard Markdown doesn't support. You might create your template such that the Agent generates a completely static file, or you can incorporate components that allows users to fill in fields of the generated document. Once you have written the template file, you need to upload the file to your project: 1. Open the project where you want to upload the template. 2. In the **Project** tab of the left sidebar, click **Add Entity**. 3. Click **Template**. Your file browser opens. 4. Select the Markdown file to upload. 5. Click **Open**. To create a custom template, you must upload a file from your local file system. You cannot create a blank template from the project editor. ## Enabling the Documentation Agent The Documentation Agent runs in a dedicated project mode and requires the Transform Agent to be disabled. Projects can operate in either Transform Agent mode or Harmonization/Documentation mode, but not both simultaneously. To use the Documentation Agent: 1. Disable the v4 Agent for the team that owns the project by opening **Metadata β†’ Teams β†’ Select team β†’ Settings β†’ Advanced**. 2. Click the **Enable Transform Agent** toggle. 3. Return to your project. Only one agent mode can be active per team at a time. # Exploration Source: https://docs.prophecy.ai/data-analysis/ai/agent/explore Find and preview data sources using the Prophecy Agent One way to leverage the Prophecy Agent is to search your SQL warehouse, explore datasets, and generate insights with simple prompts. This allows you to add the appropriate sources to your pipeline that will undergo data processing. You must have source data in your pipeline to start building transformations. The following sections describe the ways you can interact with the Agent for data exploration. ## Find tables in your SQL warehouse Ask the Agent to search for tables in your [primary SQL warehouse](/data-analysis/environment/fabrics/prophecy-fabrics). Based on your query, it returns: * A short list of relevant datasets and their descriptions. * A full list of datasets that match your criteria. You can add a dataset to the pipeline as a [Table gem](/data-analysis/gems/source-target/source-target) directly from the chat. ### Explore datasets To learn more about a dataset, you can click on it inside the chat. This opens a dialog where you can: 1. Find the location of the dataset in the warehouse. 2. Look over the schema of the dataset. 3. Preview a sample of the data in the dataset. 4. Review the data profile of the sample. 5. Open **Explore** to chat with an Agent that limits responses to the context of the dataset. Preview dataset UI This helps you validate that your data is correct without having to manually browse through your data catalogs. ## Describe a dataset If you want a summary of dataset information, you can ask the Agent to describe a table. The Agent will provide a quick overview of key metadata, including the database and schema the table belongs to, as well as the names and data types of each column. This helps you understand the structure of the dataset without having to open or query it directly. ## Compare datasets If you ask the Agent to compare datasets, you can quickly assess which dataset is more suitable as a source for your pipeline by analyzing differences in schema structure, column names, data types, and size. This allows you to identify which dataset aligns better with your pipeline's requirements, such as having the right fields, consistent naming conventions, or expected formats, without needing to inspect the full data. ## View sample rows from a table To preview data from a table, ask the Agent to return a sample. You can request a random sample from the table or specific rows, such as "the ten most recent purchases over \$100". All queries are executed under your existing warehouse permissions. The Agent returns: * A table showing the sample data. * A **Preview** option for a closer look. * SQL execution logs for transparency. If the dataset isn't already part of your pipeline, you'll also see an option to add it as a Table gem on the canvas. Data sample response ### Table preview Click **Preview** to open a larger view of the data sample. In this view, you can: * Download the data as a JSON, Excel, or CSV file. * Show or hide columns. * Add the full dataset to the canvas (if not present already). ## Visualize table data You can also ask the Agent to generate charts for data visualization. The Agent returns: * An embedded chart directly in the chat. * A Preview button to open a detailed view. * SQL execution logs showing how the chart was generated. If the chart is based on a dataset that isn't yet in your pipeline, you'll see an option to add it as a Table gem. Chart generation response ### Chart preview To see a larger version of the chart, click **Preview**. This opens the data visualization dialog, which has two tabs. | Tab | Available actions | | ------------- | ------------------------------------------------------------------------------------------------------------------------------- | | Visualization |
  • View a larger version of the chart
  • Download the chart as an image
  • Copy the chart as an image
| | Data |
  • View the underlying data
  • Download the data as a JSON, Excel, or CSV file
  • Show or hide columns
| To learn more about data visualization, see [Charts](/data-analysis/development/runs/data-explorer/charts/charts). ## Sample prompts Here are some sample prompts that you can ask to search, explore, and learn about the data. | Scenario | Prompt | | ---------------- | --------------------------------------------------------------------- | | Find dataset | "Find the dataset that shows employee hiring information and history" | | View data sample | "Return the top ten highest sales from `@daily_orders`" | | Describe dataset | "Give me more details about `@revenue_opportunities`" | | Visualize data | "Plot the sales by country" | ## Troubleshooting The Agent depends on a [knowledge graph](/data-analysis/ai/knowledge-graph/knowledge-graph) to retrieve metadata about components like datasets. If the Agent does not recognize a table that you try to reference, it could be because the knowledge graph has not been indexed recently enough to capture the dataset. For detailed instructions on indexing, see [Knowledge graph indexer](/data-analysis/ai/knowledge-graph/indexer). # Harmonization Source: https://docs.prophecy.ai/data-analysis/ai/agent/harmonization/overview Automate data standardization with harmonization Data harmonization transforms disparate source data into a standardized target Common Data Model (CDM). Source systems often use different naming conventions, data types, and structures to represent the same type of data. For example, two loan tapes might record the same value β€” say, original principal balance β€” as `OrigBal` in one system and `Original_Balance_Amt` in another, with different date formats and rounding conventions layered on top. To merge this data into a single dataset, you need to harmonize it. The Harmonization Agent does this for you by transforming source columns into their expected target columns while validating data quality. You start by selecting a source dataset, reviewing its schema, choosing a target CDM, and letting the Agent generate an initial mapping. After the initial mapping is complete, you can review mapped and unmapped columns, resolve data quality issues, preview the output, and export the harmonized data. Behind the scenes, Prophecy creates a pipeline that performs the harmonization process. ## How does the Harmonization Agent work? The Harmonization Agent uses AI to infer mappings between source schemas and a target [Common Data Model](#common-data-model) that you define. ### Common Data Model A Common Data Model (CDM) defines a target schema for source data. It specifies standardized column names, data types, and data tests used across downstream pipelines. The CDM is typically created once and reused across multiple harmonization workflows. ## Access harmonization The harmonization feature is currently available only through [the Structured Finance Edition](/data-analysis/getting-started/structured-finance-edition). To access Structured Finance, go to `https://app.prophecy.ai/finance/`. If you are already signed up for [Professional Edition](/data-analysis/getting-started/professional-edition), you will need to create a separate sign-in for Structured Finance. ## Create a new data mapping To create a data mapping: 1. Open Structured Finance. 2. Click **Tape Cracking**. 3. Select a source dataset. Either: * Click **Table** to upload an Excel, CSV, Parquet, or similar file or * Click **Select Existing** to choose an existing table. 4. Click **Continue**. ## Review the source data The source data review opens with the schema for your chosen source (that is, a page that displays column names, data types, and any available metadata). To view sample data by column, click **Data** in the upper-right corner of the page. After reviewing the schema or sample data, continue to the target data model selection. ## Select a target data model Available data models appear in the left-hand column. 1. Select a target data model. 2. Review the target model schema. 3. Click **Map Data**. Each Structured Finance project can contain one Common Data Model (CDM). To create a new CDM, either create a new project or ask the Agent to create one for you. To create a CDM manually: 1. Create a new project by clicking the **Projects** folder icon and selecting **Create New**. 2. Select **Tape Cracking**. 3. At the top of the page, click **+** > **New** > **CDMs**. 4. In the **New Common Data Model** dialog, enter a name for the CDM and click **Create**. 5. Define the schema for the CDM using one of the following methods: * Click **Add Table** to manually define tables and columns. * Click **Upload Schema** to import a schema from a CSV or other supported file. ## Generate mappings Initially, Prophecy runs a deterministic, rule-based mapping pass. The deterministic pass considers: * Exact-name matches with compatible data types * Historical, approved mappings stored in a per-CDM memory store β€” mappings previously approved for this CDM's other pipelines can be suggested again, with a distinct historical-match attribution The Agent then generates additional mapping suggestions. After processing finishes, the mapping summary shows how many target columns were mapped, how many were not mapped because corresponding source data was unavailable, and whether unresolved data quality issues remain. ## Review mapping details 1. Open the new data mapping. 2. Select **View Details**. 3. Select a column from a mapping category. 4. Review the line that displays the source-to-target mapping. 5. Select the target column to view its details. 6. Click **Next** to continue. harmonization result ## Review data quality results Open the **Data Quality** tab to review the mapping checks. The summary displays table-level and column-level checks, including how many checks passed and failed. When a check fails, click **Go to Mapping** to review the related mapping. ## Human-in-the-loop validation While the agent handles routine mappings efficiently, human oversight catches edge cases and ensures that transformations align with business rules. Before harmonization processes, you can validate agent-generated mappings and transformations. During review, Prophecy shows you how each source column is mapped to the target schema, any transformations applied, a confidence score for each mapping, and any data quality tests applied. ## Apply a suggested fix The Agent does not directly modify the live mapping when resolving a data quality failure. Instead, it creates a separate suggested fix that includes: * A proposed transformation * A plain-language description of the proposed change The suggested fix remains separate from the live mapping until you explicitly apply it. 1. Ask the Agent for suggestions to resolve the mapping issue. 2. Review the suggested fix and its description. 3. Select the proposed option you want to use. 4. Explicitly apply the suggested fix to update the mapping. Suggested options can include a default value such as `0`, a fallback to another field, or a fallback to an expected value. ## Preview and export the output 1. Click **Output Preview** to review the harmonized results. 2. Return to the mapping when additional changes are required. 3. After the mapping completes successfully, review the success message indicating that the dataset is ready. 4. Export the harmonized data. final page harmonization ## View the generated pipeline Prophecy creates a pipeline behind the scenes to perform the harmonization. Open the generated pipeline to review the implementation and inspect the details of each gem. # Supported AI models Source: https://docs.prophecy.ai/data-analysis/ai/agent/llm-support Supported models, endpoint providers, and plan requirements for Prophecy AI Agents. Prophecy AI Agents are engineered and optimized for a single model configuration to ensure consistent reasoning quality, transformation accuracy, and reliability. This document defines supported models, allowed endpoint providers, and billing requirements for Agents, as well as explicitly unsupported configurations. ## Supported model Prophecy Agents support Anthropic Opus 4.5 only. The agent architecture, reasoning workflows, and transformation logic are tuned specifically for this model. No other base models are supported. ## Alternative Claude endpoints Applicable to the [Enterprise Edition](/data-analysis/administration/platform/editions) only. Enterprise customers may choose to use their own Anthropic-hosted model endpoint instead of the Prophecy-managed endpoint. Customer-provided Anthropic endpoints are supported through the following providers: * Anthropic Console * Amazon Bedrock * Google Vertex AI * Microsoft Foundry When using a customer-provided endpoint, only PAYG, API-token-based Anthropic models are allowed. ## Lower-tier models (strongly discouraged) Customers can use lower-cost Anthropic models, such as Haiku or Sonnet, but this is strongly discouraged. Using lower-tier models will significantly degrade agent performance, including: * Lower reasoning accuracy. * Weaker transformations. * Higher failure rates. Prophecy Agents are optimized specifically for Anthropic Opus 4.5, and using alternative models introduces substantial quality risk. ## Endpoint billing and authentication requirements Only PAYG, API-token-based Anthropic models are supported. Based on supported providers, authentication mechanisms include: * API key * API key or AWS credentials * GCP credentials * API key or Microsoft Entra ID Anthropic flat monthly or per-user plans, including Claude Code licensing, are not supported and cannot be connected to Prophecy. ## Unsupported endpoints and plans The following endpoints and plans are not supported: * Databricks model endpoints. * Anthropic flat monthly or per-user plans (for example, Claude Code licensing). Prophecy does not support configuration outside of PAYG, API-token-based Anthropic models. # Prompting best practices Source: https://docs.prophecy.ai/data-analysis/ai/agent/prompting-best-practices Guidance for working effectively with the Prophecy Agent using natural language The Prophecy Agent enables you to build and modify pipelines, analyses, and documentation using natural language. Like any agent, it performs best when you provide the Agent with clear intent, structured context, and opportunities for review. This guide outlines practical techniques for working effectively with the Transform, Harmonization, and Documentation Agents. ## Why prompting best practices matter The Agent responds to prompts by implementing concrete project changes such as generating pipelines, adding gems, or updating analyses. Clear, well-structured prompts improve: * Accuracy of generated transformations. * Consistency of pipeline logic. * Efficiency of iteration. * Alignment with business requirements. ## Be specific and explicit Describe exactly what you want the Agent to do. Instead of: > "Clean the data" Use: > Remove duplicate records based on `customer_id`, filter to orders from 2024, and replace null values in `email` with `'unknown@example.com'` When prompting the Transform Agent, specify: * Source tables or gems. * Join types and keys. * Filter conditions. * Column derivations. * Output expectations. Specific instructions reduce ambiguity and improve reliability. ## Common prompting patterns Prophecy's agents supports common data workflow actions, such as: * Finding and previewing datasets. * Filtering and cleaning records. * Joining and transforming tables. * Parsing semi-structured data. * Aggregating results. * Generating visualizations. * Saving outputs as tables. Thinking in these discrete actions can help you structure clear, effective prompts. ## Build pipelines incrementally For multi-stage workflows, we recommend generating transformations step by step instead of all at once. Incremental development produces more reliable results and makes each stage easier to validate, refine, and debug. For example, a structured approach might look like: 1. Add source tables to the canvas. 2. Join datasets. 3. Apply filters. 4. Derive new columns. 5. Aggregate results. 6. Save the output as a table. After each step: * Use **Inspect** to review modified gems (these gems are highlighted in yellow). * Validate both input and output data. * Confirm the transformation matches your intent. ## Use @ mentions to reference project assets Use the `@` symbol to reference specific tables or gems in your project. See [@ mentions](/data-analysis/ai/agent/chat/mentions) for more details. Examples: > Aggregate the `@orders_cleaned` table by month. > Add a filter after `@last_join` to include only records where `region = 'West'`. > Show a sample of `@daily_sales`. Referencing project assets explicitly helps the Agent identify the correct objects and reduces unintended modifications. ## Define the expected output If you need structured output, state that explicitly. Examples: > Return only the SQL query. > Save the final result as a table. > Provide a summary of the pipeline in bullet points. > Generate documentation for the current pipeline. Clear output expectations reduce follow-up clarification and rework. The Agent can also generate visualizations when prompted clearly, as in > Create a bar chart of monthly sales from @orders\_cleaned. ## Use descriptive column and gem names Descriptive names improve the Agent's ability to generate accurate logic. If source data contains unclear labels, consider renaming columns early in the pipeline to improve downstream suggestions. Use names such as: * `customer_email` * `order_date` * `total_amount` Avoid ambiguous names such as: * `col1` * `field_a` * `x` ## Inspect, validate, and restore The Agent assists development, but you should always validate results. After each generated change: * Review modified gems using **Inspect**. * Examine input and output datasets. * Confirm logic aligns with business requirements. * Use **Restore** to revert to a previous state if needed. All changes remain versioned and visible in project history. ## Tailor prompts to the active agent Prompting best practices apply across all Prophecy Agents, with slight variations. * **Transform Agent** β€” Focus on explicit transformation logic and step-by-step pipeline construction. * **Harmonization Agent** β€” Clearly define the target Common Data Model (CDM) and review generated source-to-target mappings and data quality tests. * **Documentation Agent** β€” Specify the scope (pipeline, analysis, or project) and desired level of detail for generated documentation. Understanding the active agent mode helps you structure prompts effectively. ## Summary Effective prompting combines clarity, structure, and inspection. By providing explicit instructions, building incrementally, referencing project assets, and validating results, you can maximize the reliability and efficiency of the Prophecy Agent. # Transformation Source: https://docs.prophecy.ai/data-analysis/ai/agent/transform Transform data sources using the Prophecy Agent You can use the Prophecy Agent to generate transformations based on natural language prompts. When you describe a data operation, the Transform Agent generates the corresponding gems in the pipeline and summarizes the applied changes in the chat interface. The following sections describe how to add transformations, review modifications, and restore previous versions of the pipeline. The Transform Agent runs in the default agent mode for a project. If your team has enabled Harmonization mode, the Transform Agent must be re-enabled in team settings before transformations can be generated. ## What the Transform Agent can modify The Transform Agent operates strictly within the active [project](/data-analysis/development/projects/create-project). It can: * Create and refactor pipelines. * Generate and update analyses. * Modify datasets and related project artifacts. * Update project documentation. All changes remain fully visible and editable in the project editor. ## Execution model The Transform Agent can retrieve metadata and limited data samples to validate generated transformations. When executing SQL or running pipelines, the Agent operates under the permissions of the user who invoked it. When generating output tables or executing write operations, changes take effect immediately in the connected warehouse under the invoking user’s permissions. Existing warehouse access controls and governance policies are always enforced. ## Model architecture The Transform Agent is built on a specialized deployment of Claude Code and is optimized for Anthropic Opus 4.5. Prophecy augments model reasoning with structured tool access to project artifacts, metadata, and connected warehouse systems. The model is not used in isolation; all actions are mediated through Prophecy’s execution layer and project boundaries. Enterprise customers may configure approved Anthropic endpoints in supported deployment environments. Using alternative models may affect output quality. ## Prerequisites You need at least one [Source gem](/data-analysis/gems/source-target/source-target) in your pipeline to add transformation gems with AI chat. ## Provide a transformation To generate a transformation, enter a prompt that describes the desired data operation. The Agent returns: * One or more gems on the pipeline canvas. * A description of the applied changes. * Options to restore, inspect, or preview changes. * A group of SQL execution logs. Agent SQL logs ## Inspect pipeline changes To understand the Agent changes: 1. Select **Inspect** on the chat that generates a transformation. 2. Review the configuration panel beginning with the first modified gem. Modified gems appear in yellow. 3. Hover over the **Previous** and **Next** button to display a minimap of the pipeline. This shows you the specific gem you are viewing in the context of the pipeline. 4. Use the **Previous** and **Next** controls to move through other modified gems in sequence. 5. Examine both the input and output of each gem to confirm that the transformation produces the expected result. ## Restore a previous state of the pipeline To revert changes or try another transformation from a previous state, select **Restore** from the reply you want to revert to in the chat history. The pipeline will match the earlier version. You can also manage versions from the main project [version history](/data-analysis/development/versioning/version-control). ## Create output tables After adding various data transformations in your pipeline, ask the Agent to save the result as a table. The Agent writes the output to the default database and schema configured for your connected fabric. This allows you to persist results and reuse them in downstream workflows. If you have multiple pipelines or pipeline branches that do not terminate with tables, the Agent will not be able to create an output table. ## Project history All changes made by the Agent are saved in the project [version history](/data-analysis/development/versioning/version-control). Commits are clearly marked as authored by the Agent. If you did not save your project before interacting with the Agent, Prophecy will automatically save your changes before the Agent proceeds. ## Sample prompts Here are some sample prompts that can produce transformations in your pipeline. | Scenario | Prompt | Expected output | | ------------------ | --------------------------------------------------------- | ----------------- | | Filter records | "Filter to only include customers from California" | Filter gem | | Add transformation | "Calculate the total order value as `quantity * price`" | Reformat gem | | Clean data | "Remove rows where email is null" | DataCleansing gem | | Aggregate data | "Group by region and calculate average sales" | Aggregate gem | | Rename columns | "Rename `cust_id` to `customer_id` and `amt` to `amount`" | Reformat gem | | Save output tables | "Show me and save the final output of the pipeline" | Table gem | # Inspect Agent results Source: https://docs.prophecy.ai/data-analysis/ai/inspect-results Review and verify how the Agent produces answers When you build pipelines with the [agent](/data-analysis/ai/agent/agent), much of the structure and logic is generated for you. Your role shifts from manually assembling each step to **inspecting and verifying** that each transformation behaves as expected. You can do this directly in the visual interface by opening gems and reviewing the data they produce. ## How inspection works Inspecting a pipeline typically involves the following: * Reviewing the pipeline as a whole. * Reviewing the logic inside a gem. * Reviewing the data output after that gem runs. * (Optionally) Examining the SQL behind a gem. Inspection steps Prophecy provides a few tools to support this workflow: * View [pipelines](/data-analysis/development/pipelines/data-analysis-pipelines) on the canvas to see how data is flowing. * Open a [gem](/data-analysis/gems/gems) to inspect and modify its configuration. * Use the [Data Explorer](/data-analysis/development/runs/data-explorer/data-explorer) to view the output after a step. * Use the [Data Profile](/data-analysis/development/runs/data-explorer/data-profile) to analyze distributions, nulls, and data quality. * Use the [visual expression builder](/data-analysis/gems/visual-expression-builder/visual-expression-builder) to review or refine a gem's expression logic. * Inspect SQL using [code view](/data-analysis/gems/visual-expression-builder/visual-expression-builder#code-view). ## Inspect the overall pipeline Begin by reviewing the pipeline as a whole to understand how data flows through each step. On the pipeline canvas, look at: * **The sequence of steps** β€” how data moves from sources to final outputs. * **Joins and branching paths** β€” where datasets are combined or split. * **Inputs and outputs** β€” what each gem receives and produces. * **Key transformation points** β€” where filtering, aggregation, or enrichment occurs. This high-level view helps you answer questions like: * Does the pipeline follow the intended logic from start to finish? * Are joins happening in the right place and in the right order? * Are there unnecessary or missing steps? * Does the final output reflect the transformations you expect? Here are some examples of gems featured in the screenshot below: 1. [Table gems](/data-analysis/gems/source-target/table/snowflake), which visualize your data sources. 2. [Reformat gems](/data-analysis/gems/prepare/reformat), which transforms one or more column names or values by using expressions and/or functions. 3. [Join gems](/data-analysis/gems/join-split/join), which combine data from two or more datasets based on a shared column value. 4. [Filter gems](/data-analysis/gems/prepare/filter), which filter output rows based on defined criteria. Generated pipeline ## Inspect a pipeline step To understand what a step is doing: 1. Identify the gem you want to inspect on the canvas. 2. Open the gem to review its configuration. 3. Check how inputs are transformed: * joins and join conditions * filters and conditions * derived or reformatted columns 4. If the gem includes expressions, review them using the visual expression builder. Opening a gem helps you understand how the transformation is defined. For example, the screenshot below depicts a Filter gem that filters out rows where `MONTHLY_PREMIUM` is null or zero. Generated pipeline ## Inspect results After reviewing the logic, validate the output: 1. Run the pipeline up to and including the gem. 2. Open the result in the [Data Explorer](/data-analysis/development/runs/data-explorer/data-explorer). 3. Review: * rows and sample values * column names and types * schema changes from previous steps If you need deeper insight into the data, use the [Data Profile](/data-analysis/development/runs/data-explorer/data-profile) to check distributions, null values, and outliers. Inspecting results helps you confirm that the transformation behaves as expected. Data explorer with data analysis open ## What to look for When inspecting a pipeline step, focus on whether the output matches your expectations: * Did the number of rows change in the way you expected? * Were any rows unexpectedly dropped or duplicated? * Did the schema change correctly (new columns, renamed fields, type changes)? * Are there unexpected null values or missing data? * Do derived columns contain the correct values? * Do joins and filters behave as intended? If something looks off, return to the gem, adjust the logic, and re-run the pipeline to verify the change. ## Inspect expressions Some transformations rely on expressions defined inside a gem. Use the [visual expression builder](/data-analysis/gems/visual-expression-builder/visual-expression-builder) to: * Review comparison logic and conditions * Combine conditions with logical operators like `AND` and `OR` * Work with parameters and dynamic values After updating an expression, run the pipeline and verify the output in Data Explorer to ensure the expression produces the expected results. ## Inspect generated SQL All pipelines compile to SQL that runs in your warehouse. You can inspect this SQL to verify exactly how transformations are executed. To do this: 1. Open a gem. 2. Switch to [code view](/data-analysis/gems/visual-expression-builder/visual-expression-builder#code-view). 3. Review the generated SQL. Inspecting SQL is useful when you want to: * Confirm how joins, filters, or aggregations are implemented * Debug unexpected results at a lower level * Validate performance or query structure ## Next steps * Learn more about working with gems in the [gems overview](/data-analysis/gems/gems) * Explore outputs using the [Data Explorer](/data-analysis/development/runs/data-explorer/data-explorer) * Analyze data quality with the [Data Profile](/data-analysis/development/runs/data-explorer/data-profile) * Build and refine logic using the [visual expression builder](/data-analysis/gems/visual-expression-builder/visual-expression-builder) # Knowledge graph indexer Source: https://docs.prophecy.ai/data-analysis/ai/knowledge-graph/indexer Configure automatic indexing and authentication for the Knowledge Graph The [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) indexer retrieves metadata from your SQL warehouse and external data storage systems to build and maintain the metadata index that powers AI features in Prophecy. Prophecy automatically indexes your data environment when you create a fabric using your default credentials. After the initial run, you can configure indexing behavior to control when and how the indexer runs. This page covers: * [Scheduling automatic indexing](#configure-automatic-indexing) * [Triggering manual runs](#manually-trigger-indexing) * [Configuring separate authentication credentials for the indexer](#add-separate-authentication-for-the-indexer) ## How indexing works The Knowledge Graph indexer processes each connection in your fabric separately. For each indexing run, the indexer: 1. Uses the credentials stored in the connection to authenticate with the external system. 2. Retrieves metadata only for databases, schemas, tables, and storage locations that the configured identity is authorized to access. 3. Indexes metadata such as table names, schemas, column names, data types, and available object descriptions. 4. Updates the Knowledge Graph with the latest metadata. Prophecy indexes metadata using both structured search and embedding-based retrieval techniques to support AI-powered dataset discovery and contextual assistance. The Knowledge Graph stores metadata only. It does not store actual warehouse data values, sampled records, or summaries of table contents. ## Knowledge graph dependency checklist Before configuring or troubleshooting the indexer, verify that the following requirements are satisfied. | # | Item | Why it matters | Required? | | :-: | ----------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | --------------------- | | 1 | KG indexer has run successfully | AI features rely on indexed metadata for dataset discovery and schema-aware assistance | Required | | 2 | Knowledge Graph services are running and reachable | AI features require access to Knowledge Graph retrieval services | Required | | 3 | `KNOWLEDGE_GRAPH_BASE_URL` configured in `sql-sandbox-config-map` | Enables AI services to access Knowledge Graph retrieval APIs | Required | | 4 | `CLAUDE_MODEL` configured in `sql-sandbox-config-map` | Specifies the Anthropic model used for Knowledge Graph-powered AI assistance | Required | | 5 | `AI_DATA_ACCESS_CLUSTER_ENABLED` flag state is known | Determines whether the AI agent can perform live data operations or operate in metadata-only mode | Required | | 6 | Sufficient worker resources are available | Indexing operations require adequate compute and memory resources | Required for indexing | | 7 | KG indexing completed with SUCCESS status | Schema and dataset lookups depend on up-to-date metadata | Required | | 8 | Incremental indexing schedule confirmed | Keeps metadata current as warehouse schemas evolve | Recommended | Items 3 and 4 require updating the `sql-sandbox-config-map`. See [Knowledge graph configuration](/administration/management/cluster-admin-settings/knowledge-graph-config) for setup instructions. ## Configure automatic indexing You can configure scheduled indexing to keep your Knowledge Graph up to date without manual intervention. 1. In Prophecy, open **Metadata > Fabrics**. 2. Select the fabric where you wish to enable indexing. 3. Open the **Connections** tab. 4. Open the connection to be indexed. 5. In the connection dialog, scroll to the **Knowledge Graph Indexer** tile and enable **Knowledge Graph Periodic Indexing**. 6. Configure the schedule to run hourly, daily, or weekly. The schedule must have a defined frequency and timezone. By default, Prophecy uses the timezone from where you access the application. ### Scheduling parameters | Schedule type | Parameter | Description | Default | | ------------- | --------------------- | --------------------------------------------------------------------------------------------------------------------------- | ------------------------------------- | | Hourly | Repeat every ... from | The interval in hours between indexing runs, starting at a specific time.
Example: Repeat every 2 hours from 12:00 AM. | Every 1 hour
starting at 2:00 AM | | Daily | Repeat at | The time of day when indexing runs.

Example: Repeat at 9:00 AM. | 2:00 AM | | Weekly | Repeat on | The day(s) of the week when indexing runs.

Example: Repeat on Monday, Wednesday, Friday. | Sunday | | Weekly | Repeat at | The time of day when indexing runs.

Example: Repeat at 9:00 AM. | 2:00 AM | ## Manually trigger indexing You may need to manually trigger indexing if newly created tables or schemas are not yet available to AI features. To manually trigger indexing: 1. In Prophecy, open **Metadata > Fabrics**. 2. Select the fabric you want to index. 3. Open the **Connections** tab. 4. Open the connection to be indexed. 5. Scroll to the **Knowledge Graph Indexing Status** tile in the connection dialog. 6. Click **Start** to begin indexing and monitor progress. You can also trigger indexing from the [Environment tab](/data-analysis/environment/connections/connections#environment-browser) in your project: 1. Open a project in the project editor. 2. Attach the fabric you want to index. 3. Open the **Environment** tab in the left sidebar. 4. Locate the **Missing Tables?** callout below your connections. 5. Click **Refresh**. Prophecy may prompt you to manually trigger indexing if the AI agent cannot locate a table or schema during a conversation. ## Add separate authentication for the indexer In some environments, administrators may want more granular control over which metadata is indexed into the Knowledge Graph. For Databricks connections, Prophecy supports configuring separate authentication credentials specifically for the Knowledge Graph indexer. There are two types of credentials stored in a connection: * **Pipeline Development and Scheduled Execution credentials** control how pipelines authenticate when they run. * **Knowledge Graph Indexer credentials** control how the indexer authenticates when retrieving metadata on a schedule. If separate credentials are not configured, the indexer uses the pipeline development credentials. The Knowledge Graph indexer always uses the same identity as the pipeline development identity if the pipeline development authentication strategy is [Personal Access Token](/data-analysis/environment/connections/databricks#personal-access-token-pat) (rather than OAuth). This section does not apply when using PAT authentication. ### Prerequisites Before configuring dedicated credentials for the Knowledge Graph indexer, you must: * Upgrade to Prophecy 4.2.2 or later. * Configure your SQL warehouse connection with a [Databricks connection](/data-analysis/environment/connections/databricks). Other SQL warehouses are not currently supported for separate indexer authentication. * Be a Prophecy administrator. * Be a Databricks administrator with permission to assign appropriate access to the indexing identity. The configured identity must have sufficient permissions to retrieve metadata for the warehouse objects you want indexed into the Knowledge Graph. Knowledge Graph indexing permissions should generally match or exceed the permissions used for pipeline execution. This helps ensure that metadata for the datasets used in pipelines is also available to AI features. Prophecy does not enforce this automatically. ### Procedure To configure separate authentication for the Knowledge Graph indexer: 1. In Prophecy, navigate to **Metadata > Fabrics**. 2. Select the target fabric. 3. Open the **Connections** tab. 4. Edit the **SQL Warehouse Connection**. 5. Scroll to the **Knowledge Graph Indexer** section. 6. Configure authentication based on your pipeline development authentication method: * If you use User OAuth for \*\*pipeline development, choose either OAuth (User) or OAuth (Service Principal) for the Knowledge Graph indexer. * If you use Service Principal OAuth for \*\*pipeline development, you can only use Service Principal OAuth for the Knowledge Graph indexer. #### Service Principal OAuth (recommended) Recommended for production and scheduled indexing because credentials do not expire. * **Configuration**: Reuse pipeline development credentials or provide a separate Service Principal Client ID and Client Secret. * **Indexed metadata**: Metadata for all warehouse objects the service principal is authorized to access. If pipeline development uses **User OAuth**, Prophecy continues to enforce user-level permissions even when the Knowledge Graph indexer uses service principal credentials. #### User OAuth Recommended primarily for development environments. * **Configuration**: Uses the same app registration as pipeline development. * **Indexed metadata**: Metadata for warehouse objects the authenticated user is authorized to access. * **Limitations**: Scheduled indexing can fail when user credentials expire or require reauthentication. # What is a knowledge graph? Source: https://docs.prophecy.ai/data-analysis/ai/knowledge-graph/knowledge-graph Prophecy creates a metadata index to power AI features Prophecy uses **Knowledge Graphs** to help AI features such as Prophecy Agent and Copilot understand your [data environment](/data-analysis/environment/fabrics/prophecy-fabrics). A Knowledge Graph is a metadata index that represents relationships between schemas, tables, columns, data types, and other objects in your fabric. The Knowledge Graph contains metadata about your data environment β€” not the actual data stored in your warehouse. A Knowledge Graph is a metadata index that represents relationships between schemas, tables, columns, data types, and other objects in your fabric. For general background on knowledge graphs, see the [Wikipedia article on knowledge graphs](https://en.wikipedia.org/wiki/Knowledge_graph). The Knowledge Graph only indexes metadata that the configured warehouse identity is authorized to access. When you interact with AI in Prophecy, the platform uses the Knowledge Graph to enrich your prompts with metadata context about accessible datasets, schemas, tables, and columns. Prophecy indexes this metadata using both structured search and embedding-based retrieval techniques to help the model identify relevant schema context from natural language prompts and generate more accurate responses and SQL. ## Knowledge graph generation Each fabric maintains its own Knowledge Graph index of accessible metadata. The indexer retrieves metadata from all data connections attached to the fabric, using either your identity or a separately configured identity. During indexing, Prophecy retrieves metadata only for the warehouse objects that the configured identity is authorized to access. You can schedule automatic refreshes or trigger manual indexing to keep the Knowledge Graph current. Prophecy uses the Knowledge Graph associated with the fabric attached to your project. If you attach to a different fabric, AI features will use that fabric's Knowledge Graph. ## Agent behavior when data access is disabled The `AI_DATA_ACCESS_CLUSTER_ENABLED` flag controls whether the V4 agent can access live data. When this flag is set to `false`, the agent operates in a metadata-only mode and is limited to Knowledge Graph operations. In this mode, the agent can still use Knowledge Graph metadata to search datasets, retrieve schema information, and provide metadata-aware assistance. However, the AI agent cannot perform live data operations such as executing SQL queries, running pipelines, or generating summaries that require warehouse access. | Capability | Allowed | Agent behavior | | --------------------------------- | ------- | ------------------------------------------------------------------------------------------------------- | | KG search | βœ“ | Returns KG-indexed metadata results | | Dataset discovery | βœ“ | Lists datasets available in the KG | | Schema lookup (via KG) | βœ“ | Returns schema information stored in the KG without querying the source | | SQL queries | βœ— | Explains that live data access must be enabled | | Pipeline execution | βœ— | Explains that live data access must be enabled | | Summarization requiring live data | βœ— | Provides a partial response using available metadata and explains that live data access must be enabled | When data access is disabled, the agent can still provide metadata-aware assistance using the Knowledge Graph, but live data operations require data access to be enabled. ## What's next Configure the [knowledge graph indexer](/data-analysis/ai/knowledge-graph/indexer) to schedule automatic indexing or set up separate authentication credentials. # Add and use skills Source: https://docs.prophecy.ai/data-analysis/ai/using-skills Contribute skills to Prophecy's AI Agent for specific tasks Skills tell [Prophecy's AI Agent](/data-analysis/ai/agent/agent) how to handle a specific task: when to activate, what to read first, how to proceed, and what to hand off when done. This page covers how to add one or more skills. Skills are an iterative artifact. The structure described here reflects patterns that work well in practice, but there is no single correct form. Expect to revise a skill as you learn how the Agent interprets it and how users actually phrase their requests. You can also [use the Agent](#use-agent-to-improve-skill) to generate or refine skills. ## What a skill is A skill is a named set of instructions the Agent loads when a user request matches its description. Skills conform to the [Agent Skills open specification](https://agentskills.io/specification), which defines a portable, interoperable format for agent instructions. In Prophecy, skills live at `skills//SKILL.md` in your project. Once created, skills appear in your project alongside other project entities. In Prophecy for Business, you can find them in the [Project Browser](/data-analysis/development/studio/studio#project-browser); in [Professional Edition](/data-analysis/getting-started/professional-edition), they appear in the Browse Project panel on the right side of the workspace. ## How the agent loads skills The [Agent](/data-analysis/ai/agent/agent) reads the name and description of every installed skill on every chat. It only loads the full skill content when it determines the skill is relevant to the current task. This means: * You can maintain many skills without degrading agent output. * The description is doing active routing work, not just documentation. * A vague description means the skill may never load, or may load at the wrong time. ## Add a skill ### Add a single skill 1. From the bottom of the Project Explorer, select **Add entity > Skill**. add entity button 2. In the **Add Skill** dialog, enter a skill name. By default the skill is created in the `/skills` folder. 3. Click **Create**. The skill editor opens. 4. Enter a [description](#description) (1) and [content](#content) (2). Changes save automatically. add skill prophecy for business ### Install skills from a zip file You can bulk-install a set of skills by uploading a zip file directly to the Agent. 1. In the Agent chat, attach your zip file using **Attach** or by dragging it into the chat. 2. Enter the prompt: `Install the skills in this zip`. 3. The Agent installs all skills it finds and confirms which were added. Installed skills appear in the `/skills` directory in your project. You can invoke any installed skill by name, for example `/strats` or `/validation`. ### Add a single skill 1. At the bottom of the [Project Browser](/data-analysis/development/studio/studio#project-browser) (the rightward panel in Studio), select **Add New > Skill**. Alternatively, select **+** at the top of the tab bar, then choose **New > Skill**. 2. In the **Add Skill** dialog, enter a skill name. By default the skill is created in the `/skills` folder. 3. Click **Create**. The skill editor opens. 4. Enter a [description](#description) (1) and [content](#content)(2). Changes save automatically. add skill prophecy professional edition Bulk install from a zip file is not supported in Professional Edition. The following applies when adding a skill manually. If you installed skills from a zip file, the skill files are already written, and you can edit these directly in the skill editor. ## Skill description and content When the skill editor opens, you enter a description and content. ### Description The description is the routing signal the Agent uses to decide when to load this skill. Start with "Use when..." and be as concrete as possible. Include the user intent it targets, specific trigger phrases in quotes, any file types or contextual signals that should activate it, and where relevant, what the skill does *not* cover. A vague description gives the Agent little to work with: > Helps with data mapping. A scoped description tells it exactly when to act: > Use when the user wants to map a raw data tape to the CDM. Triggers include: 'crack this tape', 'map to CDM', 'harmonize tape'. Does not handle pipeline execution or output formatting. ### Content Here, you enter the instructions the Agent follows once the skill loads. Often, a content block starts with a one-line statement of purpose and scope: what the skill does and, where useful, what it explicitly does not do. From there, organize content around the operations the Agent needs to perform. Each operation defines a trigger condition (what the user said or did), a sequence of steps, and any output format or response pattern the Agent should follow. Operations can also specify constraints: things the Agent should not do at this stage, particularly when a skill is one step in a larger workflow and scope boundaries matter. Some patterns that appear frequently in practice: * **Prerequisites** β€” a "Before anything else" section directing the Agent to read a shared config or capabilities file before acting. * **Tunable parameters** β€” numeric thresholds documented explicitly so users know what to adjust, and so the Agent can offer to expose them as pipeline variables rather than hardcoding values. * **Scope boundaries** β€” a "What this skill does NOT do" section at the end, especially when the skill sits alongside related skills that handle adjacent tasks. * **Handoff patterns** β€” explicit routing instructions telling the Agent which other skill to invoke when the user's request moves out of this skill's scope. There is no required structure for content. Use whatever sections suit the task, named and ordered to match the work. A skill that runs a multi-step pipeline might use numbered steps and prerequisite checks, whereas one that provides background context might be a few paragraphs of prose. Skill body content stays in context for the entire session once loaded. Keep it concise β€” every line is a recurring token cost. Move large reference material to supporting files in the same skill directory and link to them from `SKILL.md`. ## Use Agent to improve skill The Agent can make changes to skills once they are in the Prophecy environment. After adding skills using the steps above, you can use the following prompts to improve Prophecy skills: ### 1. Start with discovery > "Look at \[skill name] skill. How can I improve this so that it works better for Prophecy?" This triggers an analysis of the skill's current state and generates specific recommendations before the Agent makes changes. ### 2. Request data integration > "Make this skill use actual data from the workspace instead of just asking questions" Here, you ask the agent to tune the skill so that it queries tables before conversing. ### 3. Add pipeline generation > "Update the skill to offer creating pipelines when appropriate" In Prophecy, skills become more useful when they can build artifacts such as [pipelines](/data-analysis/development/pipelines/data-analysis-pipelines) and [analyses](/data-analysis/analysis/overview). ### 4. Connect to existing work > "Make this skill integrate with the \[pipeline name] pipeline I already have" Links new skills to existing pipelines for richer, context-aware behavior. ### Pattern that works well | Step | User Says | Result | | ---- | ----------------------------------------------- | ------------------------------ | | 1 | "Look at X skill, how to improve for Prophecy?" | Get analysis + recommendations | | 2 | "yes" or "do it" | Skill gets rewritten | | 3 | "now test it" or invoke the skill | Validate the improvements | ### Key transformation questions * **From advice β†’ data-driven**: "Make it search for relevant data before asking questions" * **From text β†’ artifacts**: "Have it offer to create pipelines and dashboards" * **From generic β†’ integrated**: "Connect it to existing pipelines like \[name]" * **From one-shot β†’ workflow**: "Add steps for discovery, analysis, then artifact creation" ## Further reading [Agent Skills specification](https://agentskills.io/specification): the open standard used by Claude skills. [Anthropic's The Complete Guide to Building Skills for Claude](https://resources.anthropic.com/hubfs/The-Complete-Guide-to-Building-Skill-for-Claude.pdf): offers detailed information on building skills for Claude that may be helpful in building skills for Prophecy. # Analysis components Source: https://docs.prophecy.ai/data-analysis/analysis/analysis-components Learn about the components you can use to build interactive analyses Components are the building blocks of an analysis. They define how users interact with the underlying data pipeline during runtime. Analysis components fall into three categories: * **Interactive** β€” Let you capture runtime values from users and assign them to pipeline parameters. These values override parameter defaults when the analysis runs. * **Data Integration** β€” Display or provide data used by the pipeline, including tables, charts, file uploads, and pipeline output previews. * **Content** β€” Let you add supporting context such as headings, instructions, text, and images. By combining components, you can create interactive workflows that let users enter values, execute pipelines, and review results without editing pipeline logic. ## Interactive Interactive components enable users to assign values to [pipeline parameters](/data-analysis/development/parameters/parameters). These values can influence the behavior of each pipeline run and change the output that appears in the analysis. You can use interactive components to: * Filter pipeline output. * Select execution options. * Configure runtime behavior. * Pass values into transformations and queries. If the end user leaves an interactive field blank: * Prophecy first checks for a default value in the component. * If you did not define a default value for the component, Prophecy uses the default value of the pipeline parameter. ### Text Input The user can enter any text into the field. Only string-type parameters are supported. | Setting | Description | Required | | ------------------- | ---------------------------------------------------------------------------------------- | -------- | | Configuration field | Name of the pipeline parameter to reference. | True | | Default value | Value that appears in the field by default. Users can update this value in their config. | False | | Label | Descriptive label for the text input field. | True | | Help text | Additional information displayed below the input field to guide the user. | False | | Tooltip | Tooltip providing extra context when the user hovers over the field. | False | | Is required | Whether to make the field mandatory. | False | ### Number Input The user can enter any number into the field. The number type (such as `int`, `double`, or `long`) is determined by the pipeline parameter. | Setting | Description | Required | | ------------------- | ------------------------------------------------------------------------------------------- | -------- | | Configuration field | Name of the pipeline parameter to reference. | True | | Default value | Value that appears in the field by default. Users can update this value in their config. | False | | Label | Descriptive label for the number input field. | True | | Format | Defines how the number is displayed. Options include `Standard`, `Percent`, and `Currency`. | False | | Help text | Additional information displayed below the input field to guide the user. | False | | Tooltip | Tooltip providing extra context when the user hovers over the field. | False | | Is required | Whether to make the field mandatory. | False | ### Text Area The user can enter any text in the field. Only string-type parameters are supported. | Setting | Description | Required | | ------------------- | ---------------------------------------------------------------------------------------- | -------- | | Configuration field | Name of the pipeline parameter to reference. | True | | Default value | Value that appears in the field by default. Users can update this value in their config. | False | | Label | Descriptive label for the text input field. | True | | Help text | Additional information displayed below the input field to guide the user. | False | | Tooltip | Tooltip providing extra context when the user hovers over the field. | False | | Is required | Whether to make the field mandatory. | False | The Text Area component has the same settings as the Text Input component. However, it offers a larger text input area for the user to write in. ### Dropdown The user can select a value from a predefined list. Array-type parameters are not supported. | Setting | Description | Required | | -------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- | -------- | | Configuration field | Name of the pipeline parameter to reference. | True | | Default value | Value selected in the dropdown by default. Users can select a different option in their config. | False | | Label | Descriptive label for the dropdown field. | True | | Help text | Additional guidance displayed below the dropdown. | False | | Tooltip | Tooltip providing extra context when the user hovers over the field. | False | | Options | Where you define each dropdown option. Each option should have a value to pass to the pipeline and a label. Tooltips are optional. | True | | Is required | Whether to make the field mandatory. | False | | Allow selecting multiple options | Whether the user can select multiple options from the dropdown list. | False | ### Checkbox The user can select or unselect a checkbox. Only boolean-type configurations are supported. | Setting | Description | Required | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | | Configuration field | Name of the pipeline parameter to reference. | True | | Default value | Either `True` or `False`. This determines whether the checkbox is selected by default. Users can select or unselect the checkbox in their config. | False | | Label | Descriptive label for the checkbox itself. | True | | Caption | Additional guidance displayed below the checkbox. | False | | Tooltip | Tooltip providing extra context when the user hovers over the field. | False | | Is required | Whether to make the field mandatory. | False | ### Checkbox Group The user can select or unselect a list of checkboxes. Only array-type configurations are supported. | Setting | Description | Required | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------- | | Configuration field | Name of the pipeline parameter to reference. | True | | Default value | Whether the checkboxes are selected by default. Provide this using a comma-separated list of `True` and `False` values. Users can select or unselect the checkboxes in their config. | False | | Options | Where you define each checkbox option. | True | | Is required | Whether to make the field mandatory. | False | For each option that you add to the checkbox group, you can define the: | Setting | Description | Required | | ------- | -------------------------------------------------------------------- | -------- | | Value | Value to pass to the pipeline. | True | | Label | Descriptive label for the checkbox itself. | True | | Caption | Additional guidance displayed below the checkbox. | False | | Tooltip | Tooltip providing extra context when the user hovers over the field. | False | ### Radio Group The user can select a single option from a predefined list. Array-type configurations are not supported. | Setting | Description | Required | | ------------------- | ------------------------------------------------------------------------------ | -------- | | Configuration field | Name of the pipeline parameter to reference. | True | | Default value | Radio button selected by default. Users can update this value in their config. | False | | Options | List of available choices. At least one option is required. | True | | Is required | Whether to make the field mandatory. | False | For each option that you add to the checkbox group, you can define the: | Option Setting | Description | Required | | -------------- | -------------------------------------------------------------------- | -------- | | Value | Value that to pass to the pipeline. | True | | Label | Descriptive label for the radio button itself. | True | | Caption | Additional guidance displayed below the radio button. | False | | Tooltip | Tooltip providing extra context when the user hovers over the field. | False | ### Toggle The user can enable or disable a toggle. Only boolean-type configurations are supported. | Setting | Description | Required | | ------------------- | --------------------------------------------------------------------------------------------------------------------------------- | -------- | | Configuration field | Name of the pipeline parameter to reference. | True | | Default value | Either `True` or `False`. This determines whether to toggle is on or off by default. Users can update this value in their config. | False | | Label | Descriptive label for the toggle itself. | True | | Caption | Additional guidance displayed below the toggle. | False | | Tooltip | Tooltip providing extra context when the user hovers over the field. | False | | Is required | Whether to make the field mandatory. | False | ### Date The user can enter a date from a date picker. Only date-type configurations are supported. | Setting | Description | Required | | ------------------- | --------------------------------------------------------------------------------------- | -------- | | Configuration field | Name of the pipeline parameter to reference. | True | | Default value | Date that appears in the field by default. Users can update this value in their config. | False | | Format | Defines how the date is displayed. The default format is `MMMM d, yyyy`. | True | | First day | Specifies the first day of the week in the date picker. The default is Sunday. | True | | Label | Descriptive label for the date input field. | True | | Help text | Additional guidance below the date picker. | False | | Tooltip | Tooltip providing extra context when the user hovers over the field. | False | | Is required | Whether to make the field mandatory. | False | ## Data Integration Data integration components let users view, upload, and interact with data connected to the pipeline. These components typically display pipeline output after execution or provide additional data used during runtime. You can use data integration components to: * Preview pipeline output. * Display results in charts and tables. * Upload files for pipeline processing. * Explore runtime data interactively. ### File Upload Let the user upload their own file to replace the data of a Source gem or Table gem in the pipeline. When a user uploads a file, they have to configure the file and write it to the primary SQL warehouse of the attached fabric. This is the same mechanism that the [upload file](/data-analysis/gems/source-target/table/upload-files) feature uses. | Setting | Description | Required | | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | | Source | Choose the source table that users will replace with an uploaded file. | True | | File types | Restrict the type of files that users can upload. If you do not select any checkboxes, the user can upload any type of file that Prophecy supports. | False | | Tooltip | Add a tooltip to your component to provide help or context. | False | | Is required | Select the checkbox to make the field mandatory. | False | ### Data Preview Let the user view sample data of from a [Table gem](/data-analysis/gems/source-target/source-target) or [Visualize gem](/data-analysis/gems/report/visualize) in the pipeline. | Setting | Description | Required | | ---------- | ---------------------------------------------------------------------------------------------------- | -------- | | Data table | Table or Visualize gem that points to the data that Prophecy will display in the analysis dashboard. | True | | Label | Label to describe the data preview. | True | ### Charts Display a visualization of data from a [Table gem](/data-analysis/gems/source-target/source-target) or [Visualize gem](/data-analysis/gems/report/visualize) in the pipeline. | Setting | Description | Required | | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | | Data table | Table or Visualize gem that points to the data to visualize in the analysis dashboard. | True | | Label | Label to describe the chart. | True | | Chart type | Type of chart you wish to display (for example, bar chart or line chart). | True | | Chart configuration | Configure how the chart appears. To view information about each chart configuration, visit [Charts](/data-analysis/development/runs/data-explorer/charts/charts). | True | ## Content Content components let you add supporting context to an analysis. You can use these components to: * Add instructions and descriptions. * Organize workflows with headings and sections. * Display images and reference material. * Help users understand how to interact with the analysis. ### Text Add context to your analysis. Use the Inspect tab to add formatting to text, such as heading type, bold, italics, links, and more. | Setting | Description | Required | | ------- | ----------------------------------------------- | -------- | | Content | Text to be displayed in the analysis dashboard. | True | ### Image Embed an image into your analysis. | Setting | Description | Required | | ------------ | ------------------------------------ | -------- | | Image source | How Prophecy retrieves the image. | True | | Source URL | URL of the image. | True | | Alt text | Option to add alt text to the image. | False | # Analysis settings Source: https://docs.prophecy.ai/data-analysis/analysis/analysis-settings Configure settings for your analysis Analysis settings control how your analysis behaves during development and runtime. To access the settings for an analysis: 1. Open the project that includes the analysis. 2. Click on the analysis in the project sidebar. 3. In the same middle panel that contains the analysis, click the ellipses `...` menu. 4. Click **Settings**. ## Choose Parameter Set This setting lets you choose the [parameter set](/data-analysis/development/parameters/parameter-sets) that will be used when the analysis runs. The selected parameter set applies to all analyses by default unless the end user provides their own values. If you don't define a parameter set for an analysis, Prophecy uses the active parameter set for the pipeline. # Consume analyses Source: https://docs.prophecy.ai/data-analysis/analysis/consume-analysis Run, configure, and schedule analyses Analyses let you interact with data pipelines through a configurable runtime experience. You can use analyses to: * Enter runtime values through forms and interactive inputs. * Run pipelines without editing pipeline logic. * Review results in charts and tables. * Save reusable configurations. * Schedule recurring executions. * Share operational workflows with other teams. ## Open an analysis 1. Open the **Analysis Browser** (the graph icon) in the left sidebar. 2. Select an analysis. analysis browser page The analysis opens in **Preview** mode. Preview mode lets you interact with the analysis as an end user would, including entering runtime values, executing the pipeline, and reviewing results. ## Understanding configurations Analyses support multiple **configurations**. A configuration is a named set of runtime values for an analysis. Configurations allow teams to create reusable presets for different workflows, environments, or operational scenarios. For example, an analysis might include: * A `production` configuration * A `staging` configuration * Configurations for different regions or business units Configurations appear under the analysis in the Analysis Browser. Each analysis can contain multiple configurations, which you can select, rename, schedule, or run independently. Runtime values entered in a configuration override the pipeline's default parameter values during execution. ## Create a configuration 1. In the **Analysis Browser**, hover over an analysis. 2. Click the **+** icon. 3. Enter a configuration name. 4. Click **Create**. Configuration names must be unique within the analysis. ## Manage configurations To manage a configuration: 1. Open the configuration menu. 2. Select an action: * Rename * Delete You can switch between configurations by selecting them in the Analysis Browser. ## Run an analysis When you run an analysis, Prophecy executes the underlying pipeline using the current configuration values. Analysis results render in charts, tables, and other visualization components. 1. Open a configuration. 2. Enter or update runtime values. 3. Click **Run**. 4. Review the results. For example, you might: * Enter `West` in a **Region** field * Run the analysis * Review a chart showing sales by product category for the western region Interactive components capture runtime input, while charts and tables display the resulting pipeline output. Running an analysis always executes the full underlying pipeline. Any additional transformations or tables included in the pipeline will also run, even if they are not displayed in the analysis. ## Refresh analysis results Analyses can become out of date if the underlying pipeline changes. If the pipeline definition changes after the analysis was last executed, Prophecy displays a warning indicating that the analysis should be rerun. To refresh the analysis: 1. Open the configuration. 2. Click **Run**. ## Schedule an analysis You can schedule configurations to run automatically. Schedules are attached to individual configurations, allowing the same analysis to run with different runtime values on different schedules. For example: * A `production` configuration might run hourly * A `staging` configuration might run once per day To create a schedule: 1. Open a configuration. 2. Click **Schedule**. 3. Configure: * Cron expression * Time zone * Notification settings 4. Save the schedule. ## Share analyses You can share published analyses with other teams for interactive consumption and execution. Access to analyses is controlled by Prophecy's team-based permission model. If your team owns a project, you can: * Create analyses * Edit analyses * Delete analyses * Manage configurations If a project is shared with your team, you can: * Run analyses * Use existing configurations * Review results You cannot modify the analysis structure or underlying pipeline. To learn more, see [team-based access](/data-analysis/administration/management/users/access/team-based-access). ## What's next To learn how to build analyses, see [Create an analysis](/data-analysis/analysis/create-analysis). To customize runtime behavior and visualizations, see [Analysis components](/data-analysis/analysis/analysis-components). # Create an analysis Source: https://docs.prophecy.ai/data-analysis/analysis/create-analysis Create and build an analysis dashboard using the Agent or manually An analysis is an interactive business application built on top of a data pipeline. Analyses let teams expose pipeline logic through forms, charts, and other interactive components so users can run and consume data without writing code. When a user runs an analysis, Prophecy executes the underlying pipeline using the values entered in the analysis UI. ## How analyses work An analysis acts as a UI layer on top of a data pipeline. The general workflow is: 1. You create an analysis connected to a pipeline. 2. Interactive components are bound to pipeline parameters. 3. Users enter values in the analysis UI. 4. Prophecy executes the pipeline using those values. 5. Results render in charts, tables, and other components. Analyses support multiple saved configurations. Each configuration stores a reusable set of runtime values and can be run or scheduled independently. You can enter runtime values in the analysis, which override the pipeline's default parameter values during execution. ## Edit mode and preview mode Analyses have two primary modes: * **Edit mode** β€” Build and configure the analysis layout. Add components, configure charts, and bind inputs to parameters. * **Preview mode** β€” Run the analysis as an end user would. Enter runtime values, execute the pipeline, and review the results. ## Overview Use the Agent to automatically generate a pipeline and analysis from a prompt. ## Add source data Add a dataset to your pipeline: 1. Open the **Source/Target** gem category. 2. Click **Table**. 3. Open the gem and select **+ New Table**. 4. Choose **Seed** as the type. 5. Name the dataset `home_tech_transactions`. 6. Paste your data and click **Load Data**. 7. Click **Save**. ## Generate an analysis In the Chat interface, enter a prompt: ```text theme={null} Visualize sales performance per product category in @home_tech_transactions ``` The Agent will: * Transform the data. * Create a visualization. * Generate a pipeline. * Generate an interactive analysis dashboard. ## Continue building You can continue prompting the Agent to: * Add new visualizations * Modify transformations * Expand the dashboard ## Add interactive components To add interactive components to the analysis, you'll need to embed a pipeline parameter in the pipeline. [Pipeline parameters](/data-analysis/development/parameters/parameters) enable dynamic behavior in pipelines by allowing values to be set at runtime. In this case, pipeline parameter values are set by the end user when they run the analysis. ### Create a pipeline parameter In this section, you'll define a pipeline parameter called `region`. The parameter will capture the region that the user selects, allowing the pipeline to filter transactions based on the region the user wants to see. 1. Open the pipeline tied to the analysis. 2. Click **default** in the project header to open the parameter settings. 3. Open the **Pipeline Parameters** tab. 4. In the default parameter set, click **+ Add Parameter**. 5. Name the parameter `region`. 6. Set the parameter type to `String`. 7. Click **Select expression > Value**. 8. Enter `North` as the default value to be used during [interactive pipeline runs](/data-analysis/development/runs/execution). 9. Click **Save**. ### Add a filter Next, add a Filter gem to the pipeline. To make the filter condition dynamic, you'll use the pipeline parameter in the gem. 1. Add a **Filter** gem directly after the Source gem and before the next gem (generated by the Agent). 2. Open the **Filter** gem configuration. 3. For the filter condition: * Click **Select expression > Column** and select the `Region` column. * Click **Select operator** and select **equals**. * Click **Select expression > Configuration Variable** and select the `region` parameter. Configure visual expression to use
       parameter This expression is equivalent to `Region = {{ var('region') }}` in the Code view. 4. Click **Save**. ### Add a Text Input component Add a Text Input component to allow the end user to enter the region they want to see. 1. Open the analysis dashboard that the Agent created. 2. Click **Edit**. 3. Open the **Interactive** dropdown and select **Text Input**. 4. In the **Inspect** tab, for the **Configuration field**, select the `region` parameter. 5. For the **Label**, enter `Region`. 6. Drag the Text Input component above the bar chart visualization. For full configuration options, see [Analysis components](/data-analysis/analysis/analysis-components). ## Overview Create an analysis manually to build a dashboard from an existing pipeline. ## Add source data Add a dataset to your pipeline: 1. Open the **Source/Target** gem category. 2. Click **Table**. 3. Open the gem and select **+ New Table**. 4. Choose **Seed** as the type. 5. Name the dataset `home_tech_transactions`. 6. Paste your data and click **Load Data**. 7. Click **Save**. ## Create the analysis 1. Open the **Project Browser**. 2. Hover over a pipeline. 3. Click the **Create analysis** icon. Alternatively: * Click the **+** icon in the canvas header * Select **Create analysis** ## Configure the analysis 1. Enter a name for the analysis in **Analysis name**. Use alphanumeric characters and underscores. 2. (Optional) Enter a description in **Description**. 3. Select a pipeline in **Pipeline name**. 4. Choose a parameter set in **Parameter set**. 5. (Optional) Update the location in **Directory path**, or leave the default. 6. Click **Create analysis**. ## Build your analysis The analysis opens in **Edit** mode on the analysis canvas, where you can design the runtime experience for users. You can add three types of components: * **Interactive** β€” Capture user input * **Data Integration** β€” Display pipeline data * **Content** β€” Add text and images For full configuration options, see [Analysis components](/data-analysis/analysis/analysis-components). ## Add a chart To add a visualization: 1. Select a **chart** component. 2. Choose a **data table**. 3. Enter a **label**. 4. Select a **chart type**. 5. Configure: * Category column * Y-axis column You can further customize chart appearance, including labels and formatting. ## Add interactive components To add interactive components to the analysis, you'll need to embed a pipeline parameter in the pipeline. [Pipeline parameters](/data-analysis/development/parameters/parameters) enable dynamic behavior in pipelines by allowing values to be set at runtime. In this case, pipeline parameter values are set by the end user when they run the analysis. ### Create a pipeline parameter In this section, you'll define a pipeline parameter called `region`. The parameter will capture the region that the user selects, allowing the pipeline to filter transactions based on the region the user wants to see. 1. Open the pipeline tied to the analysis. 2. Click **default** in the project header to open the parameter settings. 3. Open the **Pipeline Parameters** tab. 4. In the default parameter set, click **+ Add Parameter**. 5. Name the parameter `region`. 6. Set the parameter type to `String`. 7. Click **Select expression > Value**. 8. Enter `North` as the default value to be used during [interactive pipeline runs](/data-analysis/development/runs/execution). 9. Click **Save**. ### Add a filter Next, add a Filter gem to the pipeline. To make the filter condition dynamic, you'll use the pipeline parameter in the gem. 1. Add a **Filter** gem directly after the Source gem and before the next gem (generated by the Agent). 2. Open the **Filter** gem configuration. 3. For the filter condition: * Click **Select expression > Column** and select the `Region` column. * Click **Select operator** and select **equals**. * Click **Select expression > Configuration Variable** and select the `region` parameter. Configure visual expression to use
       parameter This expression is equivalent to `Region = {{ var('region') }}` in the Code view. 4. Click **Save**. ### Add a Text Input component Add a Text Input component to allow the end user to enter the region they want to see. 1. Open the analysis dashboard that the Agent created. 2. Click **Edit**. 3. Open the **Interactive** dropdown and select **Text Input**. 4. In the **Inspect** tab, for the **Configuration field**, select the `region` parameter. 5. For the **Label**, enter `Region`. 6. Drag the Text Input component above the bar chart visualization. ## Configurations Analyses support multiple saved configurations. Configurations appear under the analysis in the App Browser. Each analysis can contain multiple configurations, which can be selected, renamed, scheduled, or run independently. A configuration is a named set of runtime values for the analysis. Configurations allow teams to create reusable presets for different environments, users, or operational scenarios. For example, an analysis might include: * A `production` configuration. * A `staging` configuration. * Configurations for different regions or business units. Each configuration can be run and scheduled independently. ## Run the analysis Let's return to the analysis preview so we can test the new component. 1. From the analysis, click **Back to Preview**. 2. In the **Region** field, enter `West`. 3. Click the **Run** button. 4. Review the bar chart to see the total sales by product category for the West region. When the analysis runs, Prophecy always executes the entire underlying pipeline. This means any additional transformations, tables, or other components included in the pipeline will also run, even if they're not exposed in the dashboard. Be mindful of how you design the pipeline to ensure your dashboard triggers only the intended logic. ## Share the analysis You can share published analyses with other teams for interactive consumption and execution. Access to analyses is controlled by Prophecy's team-based permission model. If your team owns a project, you have full edit access. This means that you can build, edit, and delete analyses in the project. If a project is shared with your team, you **cannot** edit any pipeline's or analysis's structure. However, you can **run** analyses from the shared project. This ensures that your data engineering team can share pipelines they developed without exposing them to changes. To learn more, reference the documentation on [team-based access](/data-analysis/administration/management/users/access/team-based-access). ## What's next To address your specific business requirements, leverage more complex [components](/data-analysis/analysis/analysis-components) to construct robust analysis dashboards. To learn about the end user experience, see [consume analyses](/data-analysis/analysis/consume-analysis). # Analyses overview Source: https://docs.prophecy.ai/data-analysis/analysis/overview Transform raw data into actionable insights with interactive dashboards Analysis dashboards let you review and visualize data from any step of a pipeline. You can describe the insights you need to the Agent, which can generate the necessary pipeline logic and resulting analysis. You can iteratively refine the Agent's output until the results meet your requirements. For production workflows, you can integrate oversight from your data team into this process. Data engineers can optimize Agent-generated pipelines to ensure data integrity and performance. Multiple business users can then interact with the final dashboards to explore data independently within the guardrails established by the engineering team. In Prophecy's **Data Prep and Analysis** model: * **Pipelines** prepare and transform data. * **Analyses** explore, visualize, and interpret that prepared data. An analysis combines: * A data pipeline that defines the underlying logic. * Interactive components that capture user input. * Visualization components that display results. * [Parameters](/data-analysis/development/parameters/parameters) that store reusable runtime values. When a user runs an analysis, Prophecy executes the underlying pipeline using the values entered in the analysis UI. Before Prophecy 4.2.5, analysis dashboards were known as Prophecy Apps. We have renamed this feature to more strongly couple analysis with pipelines. Functionality remains largely the same; however, you can now interact with the analysis using Prophecy Agents. Any apps that you created in the past are still available as analyses. ## Use cases You can build an analysis dashboard to serve a variety of use cases. We'll highlight two mental models for building analyses here; these can help you decide how to create the analysis that best serves your requirements. You can think about creating analyses in two ways: * [Fixed dashboards](#fixed-dashboards): Always run the pipeline on the same sources. The data may change, but the pipeline is always the same. * [Interactive dashboards](#interactive-dashboards): Show results based on user input. Depending on how you build the dashboard, users can do things like upload their own data or change the parameters of the pipeline. ### Fixed dashboards Fixed dashboards display results from pipeline runs without requiring user input. They're perfect when you want to: * Preview pipeline outputs before publishing to BI tools. * Schedule regular reports with different parameter configurations. * Share a specific analysis or report with stakeholders. When you create a fixed dashboard, you include **data** and **content** components. You don't include any interactive components. ### Interactive dashboards Interactive dashboards let users customize inputs and run pipelines on demand. They're ideal when you want to: * Enable team members to answer their own questions without writing code. * Provide guardrails that guide users toward correct, tested execution paths. To make a dashboard interactive, you'll need to include **interactive** components as you build it. Interactive components let users: * Upload custom data to replace pipeline input sources. * Set pipeline parameters (dates, thresholds, regions) through simple form fields. * Save personal configurations while maintaining a shared, standardized pipeline. ## What's next? * Follow the tutorial to [create an analysis](/data-analysis/analysis/create-analysis) or review different [analysis components](/data-analysis/analysis/analysis-components). * Learn how users can [consume analyses](/data-analysis/analysis/consume-analysis) that you have shared with them. # Replays Source: https://docs.prophecy.ai/data-analysis/collaboration/project-replays Share interactive, step-by-step walkthroughs of pipeline development Replays are available in [Free and Professional Editions](/data-analysis/administration/platform/editions) only. Replays transform your pipeline development history into interactive tutorials. By breaking down the build process into step-by-step walkthroughs, replays help viewers understand not just what your pipeline does, but how and why you built it. A replay exposes fabric metadata including schemas, tables, connections, and data previews. Only share replays when this information can be safely disclosed. ## Prophecy onboarding Prophecy's onboarding experience uses replays to introduce new users to the platform. To view these replays after your initial setup, click **Onboarding** from the homepage. Reviewing the onboarding replays can help you become familiar with the tool before using it yourself. ## Build a new replay To create a replay from your pipeline's commit history: 1. Click the version control menu and select **Share Project & Replays**. 2. Navigate to the **Replays** tab. 3. Click **Build Replay** to open the Replay Builder. 4. Click **Get Started**. Prophecy automatically generates replay steps based on your pipeline's saved commits. Each step appears in the bottom panel with a default title and description. To learn more about commit history, visit [Versioning](/data-analysis/development/versioning/version-control). ### Customize your replay To refine the replay experience: 1. Click the gear icon in the bottom panel to edit: * Step order (drag to reorder) * Step titles and descriptions * Which commits to include (unselect to skip) 2. Mark steps as **Challenge Mode** to create interactive learning steps. Note that while you can reorganize steps and disable commits, you cannot change the order of commits. ### Challenge mode Challenge mode transforms passive viewing into active learning. When you mark a step as a challenge: * Viewers must recreate the pipeline transformations themselves. * The replay validates their work against the original implementation. * Viewers can switch to **Watch** mode if they need help completing the challenge. This makes replays powerful training tools for teaching data transformation techniques in Prophecy. ## Share a replay Each project can contain one or more replays. To share a replay: 1. Open a project in the project editor. 2. Open the version control menu in the top right corner. 3. Select **Share Project & Replays**. This opens the sharing dialog. 4. Click on the **Replays** tab. 5. Copy the link and send it to your audience. Anyone with the link will be able to view the replay, no login required. ## Viewer experience Users who open your replay link can: * Progress through each commit step-by-step. * See data samples automatically as the replay advances. * Pause the replay to explore the pipeline. * Attempt challenge mode steps to test their skills. * Validate their work against your original pipeline logic. ### Resuming and exiting If viewers leave a replay before finishing, they can resume from where they left off from the homepage. When viewers click **Exit** on a replay, they enter **View** mode, which is the same read-only interface used for [public sharing](/data-analysis/collaboration/project-sharing). In View mode, they can explore the completed pipeline but cannot make changes. # Project sharing Source: https://docs.prophecy.ai/data-analysis/collaboration/project-sharing Collaborate by sharing projects for development or viewing Share your Prophecy projects with team members or external stakeholders. You can grant full editing access through team membership or provide read-only access via public links. ## Share with your team Give users project access by adding them to the team that owns the project. Team members can view and edit all projects owned by their team. ### Add users from the project You can add users from a project in [Free and Professional Editions](/data-analysis/administration/platform/editions) only. To invite team members directly from your project: 1. Open the version control menu in the top right corner of the project editor. 2. Select **Share Project & Replays**. This opens the sharing dialog. 3. Navigate to the **Team Access** tab. 4. Enter one or more email addresses. 5. Click **Send Invitation**. Prophecy sends an invitation link to each email address. When clicked, Prophecy adds the recipient to the project's team. ### Add users from team settings To invite team members from a team's Metadata page: 1. Navigate to **Metadata** > **Teams**. 2. Select the team that owns your project. 3. Open the **Members** tab. 4. Click **+ Invite Users**. 5. Enter one or more email addresses. 6. Choose a role: **User** or **Admin**. 7. Click **Send Invitation**. Prophecy sends an invitation link to each email address. When clicked, Prophecy adds the recipient to the project's team. ## Share publicly Public sharing is available in [Free and Professional Editions](/data-analysis/administration/platform/editions) only. Share your project with anyone using a public link. This provides read-only access without requiring authentication, making it ideal for demos, training, or stakeholder reviews. Public links expose metadata from your fabric, including schema names, table lists, connection details, and data previews. Only share projects publicly when this information can be safely disclosed. ### Enable public sharing To create a public share link to a Prophecy project: 1. Open the version control menu in the top right corner of the project editor. 2. Select **Share Project & Replays**. This opens the sharing dialog. 3. Click on the **Public Share** tab. 4. Choose your sharing method: * Copy the link to share manually * Enter email addresses to send the link automatically To revoke public access, change the **Anyone on the internet with the link** dropdown to **No access**. ### Read-only capabilities Users with the public link can: * View pipeline structure and gem configurations. * Run pipelines and see execution results. * Inspect interim data samples between pipeline steps. * Use the Agent to ask questions about data and logic. Users cannot: * Edit pipeline components or configurations. * Use the Agent to modify the pipeline. ### Cloning shared projects When you click **Edit** while in read-only mode: 1. You are prompted to log in (if not authenticated). 2. When logged in, you will see the **Clone Project to Edit** dialog. 3. You can click **Clone Now** to create a copy of the project in your personal team. When you clone a project, the data ingress/egress gems (Source, Target, and Table) retain their original [connection](/data-analysis/environment/connections/connections) configurations from the original team's fabric. If you don't have access to that fabric, you'll need to update these gems to use connections available in your personal team's fabric instead. For example, if the original pipeline has a gem that reads from Snowflake using a connection called `Snowflake_Prod_Data`, you can either replace this connection with one from your personal fabric or request access to the original fabric. # Dependencies Source: https://docs.prophecy.ai/data-analysis/development/extensibility/dependencies Make use of external or custom components in your projects Dependencies allow you to reuse logic in your SQL projects, so you can build on work that's already been tested and versioned. Dependencies are scoped at the project level, and can include [packaged Prophecy projects](/data-analysis/development/extensibility/package-hub/package-hub), as well as external packages from GitHub or the dbt Hub. Because packages can be improved over time, you can update your project dependencies whenever a new version is published. Your project will continue to use the current version until you choose to upgrade. SQL Dependencies ## Dependency types There are three types of dependencies for SQL projects. ### Prophecy Project When you import a project from the Package Hub as a dependency, you gain access to all its components, including pipelines, gems, and functions for use in your own project. If a new version of the project is [published](/data-analysis/development/versioning/version-control), you can update your dependency version to take advantage of the latest changes. Prophecy Project dependencies have the following parameters: | Parameter | Description | | -------------------- | ------------------------------------------------------------------------------------------------------------------------- | | Project Dependencies | Choose from a dropdown list of compatible projects that are published in the Package Hub. | | Auto-upgrade | When enabled, the package will automatically upgrade to the latest version available. No manual updates will be required. | ### GitHub Dependencies can be saved to GitHub repositories and imported from there. GitHub dependencies have the following parameters: | Parameter | Description | | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Git Repository | Link to the GitHub repository containing the dependency. | | Revision | Git tag, commit hash, or branch name that points to the version that you want to import. | | Sub Directory | Path of the subdirectory in the repository that contains the dependency (if not in the root directory). | | Warn unpinned | Whether to point to your repository without specifying any version, commit, or branch. Enabling this option may result in unexpected behavior if there are changes to your latest default branch. | ### dbt Hub Packages in the [dbt Hub](https://hub.getdbt.com/) can be used to extend typical SQL functionality. If the package contains macros, you can use them via the [Macro](/data-analysis/gems/custom/macro) gem. dbt Hub dependencies have the following parameters: | Parameter | Description | | --------- | -------------------------------------------------- | | Package | Package name or path (e.g., `dbt-labs/dbt_utils`). | | Version | Package version or version range. | ## Add a dependency to your project To add a dependency to your project: 1. Open the project settings menu beside the project name in the project header. 2. Select **Dependencies**. 3. Click **+ Add Dependency**. 4. Choose the dependency type (Project, GitHub, or DBT Hub). 5. Fill in the required fields to import the correct package version. 6. Click **Create**. 7. Click **Reload and Save** to validate and download the new dependency. Dependencies are stored at the project level within the `packages.yml` file of the project code. You can also browse packaged projects in the [Package Hub](/data-analysis/development/extensibility/package-hub/package-hub) and import them from there. dbt Hub dependency ## Use dependency components in your project When you add a dependency to your project: * The dependency will be visible in the **Packages** section of the left sidebar. Expand a dependency to view the various functions, gems, and other components that it may contain. You can drag these components directly onto the visual canvas. * Any gems in the package will automatically appear in the gem drawer on the canvas itself. ## Manage your project dependencies To manage your project dependencies, there are two options: **A)** Open your project in the Studio. Then, open the project settings menu beside the project name and select **Dependencies**. **B)** Open your project in the Studio. Then, hover the **Packages** section of the left sidebar and click the **gear** icon. From here, you can add, edit, update, or remove dependencies from your project. # Functions Source: https://docs.prophecy.ai/data-analysis/development/extensibility/functions Build functions with SQL macros to be used in gem expressions In SQL projects, functions are SQL macros that transform data at the column level. Unlike gems, which operate at the table level, functions apply transformations to individual columns, making them useful for data cleansing, formatting, and complex calculations. Functions are compiled into code using the Jinja templating language (standard for dbt macros). You will use functions as [expressions](/data-analysis/gems/visual-expression-builder/visual-expression-builder) in gems. ## Create a function To add a new function to your project, perform the following steps. 1. Open a SQL project. 2. Click **Add Entity** in the project sidebar. 3. Select **Function**. 4. Name the function. 5. Keep the function in the `macros` directory. 6. Click **Create**. This opens the function configuration screen. ## Build the function You can build functions visually by populating the following fields. | Field | Description | | ----------- | -------------------------------------------------------------------------------------------------------- | | Description | A summary of what the function will do. | | Parameters | The parameters (arguments) that will be passed to the function. Parameters can be values or table names. | | Definition | SQL code that will be executed by the function. | ## Example: Concatenate columns Use the following example to learn how to build a function and use it in your pipeline. This example demonstrates a function that concatenates values from a `first_name` and `last_name` column in a customer table. 1. Click **+Add Entity** in the project sidebar. 2. Select **Function**. 3. Name the function `concat_name`. 4. Click **Create**. 5. Add the following description: `Concatenates customer first and last names in a new column`. 6. Add two parameters to the function: `first_name` and `last_name`. 7. Add the following to the macro body ```jinja theme={null} CONCAT( UPPER(LEFT({{ first_name }}, 1)), LOWER(SUBSTRING({{ first_name }}, 2)), ' ', UPPER(LEFT({{ last_name }}, 1)), LOWER(SUBSTRING({{ last_name }}, 2)) ) ``` To use this function in your pipeline: 1. Add a Reformat gem to the pipeline canvas. 2. For the target column, create a new column named `full_name`. 3. For the expression, select **Function > concat\_name**. 4. For the function parameters, choose a first name columns and a last name column. 5. Save and run the gem. In summary, this function inside of the Reformat gem let us combine customer first and last names into a new column: `full_name`. # Schema name resolution in SQL projects Source: https://docs.prophecy.ai/data-analysis/development/extensibility/generate-schema-name How Prophecy resolves schema names in SQL projects and how to restore the default macro if needed When you create a SQL project in Prophecy, the project is initialized with a small dbt macro named `generate_schema_name` under the `macros/` directory. This macro controls how Prophecy resolves the final schema name for every table in your pipeline. This page explains what the macro does, what happens if you delete it, and how to restore it. Do not delete `macros/generate_schema_name.sql` from your project. Deleting it does not reset defaults. Instead, it switches your project to dbt's concatenating default, which concatenates schema names in a way that is rarely correct for Prophecy projects. Prophecy auto-generates this macro for every new SQL project. If you started a project before this auto-generation rolled out, or you removed the file at some point, you will need to add it back manually. See [Restore the macro](#restore-the-macro). ## What happens if you delete the macro If you delete `macros/generate_schema_name.sql` from the project, dbt falls back to its built-in version of the macro. The built-in version concatenates the fabric's default schema with whatever you set on the Table gem, joined by an underscore. For example, with a Databricks fabric whose connection schema is `default` and a Table gem whose target schema is `my_schema`: | Macro present | Resolved schema | Resulting fully qualified name | | --------------------------------------------------- | ------------------- | ------------------------------------- | | Prophecy `generate_schema_name` | `my_schema` | `.my_schema.
` | | dbt built-in (`generate_schema_name` macro deleted) | `default_my_schema` | `.default_my_schema.
` | You will typically notice this in one of two ways: * The **preflight check** in the Prophecy editor reports that the target table cannot be found, because the concatenated schema does not exist in the warehouse. * A pipeline run creates a new schema named `_` in your warehouse and writes data there, instead of the schema you configured. If you only see this issue on tables that have an explicit target schema set, and tables without a target schema still resolve correctly, the missing macro is the most likely cause. ## How schemas are resolved Each Table gem in a SQL project has two schema-related inputs: * The **fabric default schema** β€” set on the Databricks, Snowflake, or BigQuery [connection](/data-analysis/environment/connections/connections) used by the fabric. dbt exposes it as `target.schema`. * The **table target schema** β€” an optional override set per Table gem. dbt passes it to the macro as `custom_schema_name`. The Prophecy `generate_schema_name` macro resolves these two inputs as follows: ```jinja macros/generate_schema_name.sql theme={null} {# Autogenerated by prophecy #} {% macro generate_schema_name(custom_schema_name=none, node=none) -%} {%- set default_schema = target.schema -%} {%- if custom_schema_name is none -%} {{ default_schema }} {%- else -%} {{ custom_schema_name | trim }} {%- endif -%} {% endmacro %} ``` The contract is straightforward: * If a Table gem **does not** set a target schema, the table is written to the fabric default schema. * If a Table gem **does** set a target schema, the table is written to that schema exactly as written. ## Restore the macro If you have already deleted `generate_schema_name`, you can recreate it from the project sidebar. Open the project in the Studio. In the left sidebar, expand the **Functions** section and confirm that no entry named `generate_schema_name` exists. If `prophecy_tmp_source` is the only function listed, the schema-resolution macro has been removed. 1. Click **+ Add Entity** in the project sidebar. 2. Select **Function**. 3. Name the function `generate_schema_name`. 4. Keep the function in the `macros` directory. 5. Click **Create**. Replace the default function body with the following content: ```jinja macros/generate_schema_name.sql theme={null} {# Autogenerated by prophecy #} {% macro generate_schema_name(custom_schema_name=none, node=none) -%} {%- set default_schema = target.schema -%} {%- if custom_schema_name is none -%} {{ default_schema }} {%- else -%} {{ custom_schema_name | trim }} {%- endif -%} {% endmacro %} ``` Save the function. Open any pipeline that writes to a Table gem with an explicit target schema. The fully qualified name in the gem preview and in the preflight check should now use the target schema directly, without the fabric default schema as a prefix. A table with target schema `my_schema` resolves to `.my_schema.
`, not `.default_my_schema.
`. ## Customizing the macro The macro is a regular dbt macro and you can customize it if you have specific routing rules. For example, you could prefix schemas by environment or by team. As long as the customized macro is named `generate_schema_name` and lives in the `macros/` directory, dbt will use it instead of the built-in version. If you customize this macro, follow the dbt guidance in [Custom schema names](https://docs.getdbt.com/docs/build/custom-schemas#how-does-dbt-generate-a-models-schema-name). Keep the macro deterministic so that the schemas you see in the editor match the schemas that the pipeline writes to at run time. ## Related topics * [Functions](/data-analysis/development/extensibility/functions) β€” Other user-built macros in SQL projects. * [Databricks connection](/data-analysis/environment/connections/databricks#connection-parameters) β€” Where the fabric default schema is configured for Databricks. * [Snowflake connection](/data-analysis/environment/connections/snowflake#connection-parameters) β€” Where the fabric default schema is configured for Snowflake. * [dbt's documentation on custom schemas](https://docs.getdbt.com/docs/build/custom-schemas) β€” Upstream documentation for the macro. # Package hub for Data Analysis Source: https://docs.prophecy.ai/data-analysis/development/extensibility/package-hub/package-hub Create and share reusable pipeline components To extend the functionality of a project, you can download **packages** from the Package Hub. Packages are versioned projects that contain shareable components, such as pipelines, gems, business rules, user-defined functions, jobs, macros, models, and more. Package Hub landing page ## Publish to the Package Hub To create reusable components for yourself and others: 1. Create a project. 2. Build the component(s). 3. Release the project (create a new project version). 4. Share your project with other teams in the Access tab of the project metadata page. 5. Publish the project to the Package Hub. Importantly, if you add a project to the Package Hub, **all of its components will be available for reuse**. Publish to Package Hub ## Access and permissions Packages in the Package Hub are only available to users in teams that you have shared the project with. After you share your dependency with a team, users in the team can add your project as a dependency to their new or existing projects. If the team does not see the package listed when they try to add it as a dependency, be sure the new project and dependent project use the same language, such as SQL, Scala or Python. For example, if the new project is a SQL project, only SQL Packages can be added as dependencies. ## Create a new package version When you update a project that is published as a package, the changes will only be available in the Package Hub when you release the project as a new version. The release must be made from the branch specified on project creation (usually `main` or `master`). This ensures that teams will review the code before releasing the project. If you want to change how the package works for a particular project without changing the original package, clone the packaged project and make your changes. ## Import a package as a project dependency There are a few different ways to add a package to a project: * Open the project and click **+** in the Gem Drawer. * Open the project dependencies and add a dependency. * Open the package in Package Hub and select **Use Package**. Import from Package Hub You cannot change package components from the target project. You can only change the components from the source project. In other words, imported components are **read-only**. ## Prophecy-provided packages Explore Prophecy-provided packages in the Package Hub to find extra gems and pipelines that may be helpful for your projects. Browse the Package Hub to explore what components may work for you. Each Prophecy-provided gem has corresponding documentation in the [Gem reference](/data-analysis/gems/gems). # Template Hub Source: https://docs.prophecy.ai/data-analysis/development/extensibility/package-hub/template-hub Play with sample pipelines provided in the Template Hub The Package Hub also includes a set of sample pipelines in the Template Hub. These pipelines cater to specific industry use cases. You can use these pipelines to learn how to run data processing flows and configure gems in a familiar context. The Template Hub is only available for users who joined Prophecy after version 4.1.0. ## Use a pipeline template At the top of the Package Hub homepage, select the **Template Hub** tab. As you browse through the available pipelines, keep in mind that each template contains several tables, and at least one pipeline. There are two ways to use a pipeline template. * **Read-only mode**: If you click on a pipeline from the Template Hub, this opens the pipeline in read-only mode. You can explore all project configurations and run any pipeline in the project, but no changes you make will be saved. * **Edit mode**: If you click **Edit** on a pipeline in the Template Hub, a copy of the project is created in your [personal team](/data-analysis/administration/management/teams/teams). This project is editable because it is completely separate from the original pipeline template. You can always open the original from the Template Hub to reference a working pipeline. Prophecy provides a Databricks compute environment for pipeline templates. This means that you don't need your own fabric to run these pipelinesβ€”you'll use the Prophecy-provided fabric. This fabric will be available automatically in the pipeline template project. ## List of pipeline templates The following lists the SQL pipelines available in each category. ### Transportation *** #### `l2_gold_airport_delay_metrics` Computes the airport-level delay metrics using flight data from the past 30 days. It aggregates departure and arrival delays, cancellation rates, and overall on-time performance for both origin and destination airports. **Gems used**: [Table](/data-analysis/gems/source-target/source-target), [Join](/data-analysis/gems/join-split/join), [Filter](/data-analysis/gems/prepare/filter), [Aggregate](/data-analysis/gems/transform/aggregate), [OrderBy](/data-analysis/gems/prepare/order-by) ### Healthcare *** #### `L0_raw_clinical_trials` Cleans up clinical trial data by filtering out malformed records, transforming nested data into tabular format, and removing leading and trailing whitespaces from column values. **Gems used**: [Table](/data-analysis/gems/source-target/source-target), [Filter](/data-analysis/gems/prepare/filter), [FlattenSchema](/data-analysis/gems/prepare/flatten-schema), [DataCleansing](/data-analysis/gems/prepare/data-cleansing) *** #### `L1_hipaa_compliance` Makes the patient ID, age, and ZIP code HIPAA compliant. **Gems used**: [Table](/data-analysis/gems/source-target/source-target), [Reformat](/data-analysis/gems/prepare/reformat) *** #### `L2_clinical_trial_gap_analysis` Performs a gap analysis to determine the cost of each diagnosis without clinical trials. **Gems used**: [Table](/data-analysis/gems/source-target/source-target), [Filter](/data-analysis/gems/prepare/filter), [Aggregate](/data-analysis/gems/transform/aggregate), [OrderBy](/data-analysis/gems/prepare/order-by), [DataCleansing](/data-analysis/gems/prepare/data-cleansing), [Join](/data-analysis/gems/join-split/join) *** #### `L2_gold_daily_billing_report` Determines the daily total billing based on patients per provider. **Gems used**: [Table](/data-analysis/gems/source-target/source-target), [Reformat](/data-analysis/gems/prepare/reformat), [Filter](/data-analysis/gems/prepare/filter), [Aggregate](/data-analysis/gems/transform/aggregate) ### Manufacturing *** #### `L2_defect_rate_analysis` Analyzes production job defect rates by calculating the defect rate per job for each production line and month, then ranks these rates into percentiles to identify how jobs compare within the same line and time period. **Gems used**: [Table](/data-analysis/gems/source-target/source-target), [Reformat](/data-analysis/gems/prepare/reformat), [WindowFunction](/data-analysis/gems/transform/window) *** #### `L2_part_processing_metrics` Calculates the average, median, minimum, and maximum for manufacturing parts by measuring the hours taken to start and complete events for each part number. **Gems used**: [Table](/data-analysis/gems/source-target/source-target), [Aggregate](/data-analysis/gems/transform/aggregate), [Join](/data-analysis/gems/join-split/join), [Reformat](/data-analysis/gems/prepare/reformat) ### Finance *** #### `L2_customer_financial_profile` Creates a detailed customer financial profile while masking sensitive information by combining raw data from various source tables and summarizing each customer's account activity, loan details, credit score history, and transaction spending. **Gems used**: [Table](/data-analysis/gems/source-target/source-target), [Reformat](/data-analysis/gems/prepare/reformat), [OrderBy](/data-analysis/gems/prepare/order-by), [Aggregate](/data-analysis/gems/transform/aggregate), [Join](/data-analysis/gems/join-split/join) ### Retail *** #### `L2_gold_campaign_analysis` Calculates campaign-level return on investment (ROI), summarizes total earnings and costs per product, and computes ROI as the percentage gain relative to marketing spend. **Gems used**: [Table](/data-analysis/gems/source-target/source-target), [Aggregate](/data-analysis/gems/transform/aggregate), [Join](/data-analysis/gems/join-split/join), [Filter](/data-analysis/gems/prepare/filter), [FlattenSchema](/data-analysis/gems/prepare/flatten-schema), [Reformat](/data-analysis/gems/prepare/reformat) ### Energy *** #### `l2_monthly_energy_report` Calculates the percentile rank of each energy asset based on its monthly energy production and benchmarks asset performance across dimensions such as state, region, and asset type. **Gems used**: [Table](/data-analysis/gems/source-target/source-target), [Reformat](/data-analysis/gems/prepare/reformat), [Aggregate](/data-analysis/gems/transform/aggregate), [WindowFunction](/data-analysis/gems/transform/window), [Join](/data-analysis/gems/join-split/join), [Reformat](/data-analysis/gems/prepare/reformat), [Union](/data-analysis/gems/join-split/union), [OrderBy](/data-analysis/gems/prepare/order-by) # SQL gem builder Source: https://docs.prophecy.ai/data-analysis/development/extensibility/sql-gem-builder Build custom gems for SQL models and pipelines Available for [Enterprise Edition](/data-analysis/administration/platform/editions) only. [Gems](/data-analysis/gems/gems) handle individual data processing tasks in a pipeline or model. While Prophecy offers dozens of gems out-of-the-box, you might want to create your own gems. The SQL Gem Builder lets you create and publish your own custom gems for SQL projects. Be sure to develop your gem code using the SQL dialect of your warehouse. ## Gem language The SQL Gem Builder supports Databricks SQL and Snowflake SQL. It's built on dbt Coreβ„’, allowing you to build upon existing dbt libraries to define new macros to use in your custom gem. You can create a gem that writes a reference to either of the following options: * A new user-defined macro * An existing macro present in a dependency (such as `dbt-utils`) ## Getting started Get started creating your own gem: 1. Open a SQL project, and the click **Add Gem**. 2. Enter a **Gem Name**. 3. Choose a **Category**. 4. Verify the **Directory Path** where your gem SQL query will be stored. 5. Click **Create**. This will open the files you will need to edit to define your custom gem. Gem builder new gem ## Steps A gem is made up of multiple components that determine the UI and logic of the gem. The Gem Builder breaks these components into steps for you while you create your gem. There are three parts to creating a gem: 1. [Create SQL Query](#create-a-sql-query) 2. [Customize Interface](#customize-the-interface) 3. [Preview](#preview-your-gem) First, you'll define the SQL query using a new or existing macro. You'll then need to customize the UI and logic of your gem. Finally, you can preview your gem. ## Create a SQL query Prophecy gems are powered by macros. Therefore, you can either define a new macro or leverage an existing one for your custom gem. Gem builder create SQL query Existing dbt macros can help define table-to-table transformations. Consider using them to complete your SQL Query. See the [dbt utils source code](https://github.com/dbt-labs/dbt-utils/tree/main/macros/sql) for macro definitions. ## Customize the interface Customizing your gem involves editing the code for specific classes, functions, and methods. Gem builder customize interface The code starts with a list of imports from the Prophecy codebase to help get you started. ```sql theme={null} from dataclasses import dataclass from collections import defaultdict from prophecy.cb.sql.Component import * from prophecy.cb.sql.MacroBuilderBase import * from prophecy.cb.ui.uispec import * ``` The following sections describe how to make edits to your gem's interface. ### Parent class Every gem class needs to extend a parent class from which it inherits the representation of the overall gem. This includes the UI and the logic. You can determine the name and category of your gem, which are `"macro_gem"` and `"Custom"` in this template. ```sql theme={null} class macro_gem(MacroSpec): name: str = "macro_gem" projectName: str = "snowflake_docs" category: str = "Custom" ``` ### Properties classes There is one class that contains a list of the properties to be made available to the user for this particular gem. Think of these as all the values a user fills out within the template of this gem, or any other UI state that you need to maintain. * A collection of input tables, represented as input ports (optional). * A configurable set of additional parameters through the dialog (optional). The content of these `Properties` classes is persisted in JSON and stored in Git. These properties can be **set** in the `dialog` function by taking input from user-controlled UI elements. The properties are then available for reading in the following functions: `validate`, `onChange`, and `apply`. ```sql theme={null} @dataclass(frozen=True) class macro_gemProperties(MacroProperties): # properties for the component with default values parameter1: str = "'default_value_of_parameter1'" ``` Additional information on these functions are available in the following sections. ### Dialog (UI) The `dialog` function contains code specific to how the gem UI should look to the user. * Automatically generated based on parameters (default). * Custom dialogs using Python or visual configurations. ```sql theme={null} def dialog(self) -> Dialog: return Dialog("Macro").addElement( ColumnsLayout(gap="1rem", height="100%") .addColumn( Ports(allowInputAddOrDelete=True), "content" ) .addColumn( StackLayout() .addElement( TextBox("Table Name") .bindPlaceholder("Configure table name") .bindProperty("parameter1") ) ) ) ``` After defining a gem in the code editor, you can preview and test it. This feature directly renders the interface for the selected gem using a dummy schema, enabling you to configure and experiment with the gem's UI components. You can then finalize them by previewing the generated SQL code. Gem builder preview There are various UI components that can be defined for custom gems such as scroll boxes, tabs, and buttons. These UI components can be grouped together in various types of panels to create a custom user experience when using the gem. After the Dialog object is defined, it's serialized as JSON, sent to the UI, and rendered there. Depending on what kind of gem is being created, a `Dialog` needs to be defined. #### Column selector You can use the column selector property if you want to select the columns from UI and then highlight the used columns using the `onChange` function. The function defines the changes that you want to apply to the gem properties once changes have been made from the UI. For example, in the reformat component provided by Prophecy, based on the columns used on the expression table `onChange` highlights the columns used on the input schema. It is recommended to try out this dialog code in gem builder UI and see how each of these elements looks in UI. ### Validation The `validate` method performs validation checks so that in the case where there's any issue with any inputs provided for the user an Error can be displayed. You can add any validation on your properties. * Optional functions such as `onChange` or `validate`, which are executed on user actions. They can dynamically alter the state of how the gem works based on the user input. ```sql theme={null} def validate(self, context: SqlContext, component: Component) -> List[Diagnostic]: # Validate the component's state return super().validate(context,component) ``` ### State changes The `onChange` method is given for the UI State transformations. You are given both the previous and the new incoming state and can merge or modify the state as needed. The properties of the gem are also accessible to this function, so functions like selecting columns, etc. are possible to add from here. ```sql theme={null} def onChange(self, context: SqlContext, oldState: Component, newState: Component) -> Component: # Handle changes in the component's state and return the new state return newState ``` ### Apply The code for invoking the macro with the gem logic is defined in the `apply` function. Here the above User Defined properties are accessible using `self.projectName.{self.name}`. ```sql theme={null} def apply(self, props: macro_gemProperties) -> str: # generate the actual macro call given the component's state resolved_macro_name = f"{self.projectName}.{self.name}" non_empty_param = ",".join([param for param in [props.parameter1] if param != '']) return f'{{{{ {resolved_macro_name}({non_empty_param}) }}}}' ``` ### Macro properties When Prophecy parses a macro invocation, it represents a macro definition in a default state. `MacroProperties` consists of the following: * macro name * project name * parameters used For example, if macro invocation is ```sql theme={null} dbt_utils.deduplicate(relation, partition_by, order_by)` ``` then Prophecy parses it into an object such as the following: ```sql theme={null} MacroParameter(value="relation"), MacroParameter(value="partition_by"), MacroParameter(value="order_by") ``` This object now has to be converted into the gem state defined by the user. This logic is defined in `loadProperties`. ```sql theme={null} def loadProperties(self, properties: MacroProperties) -> PropertiesType: # load the component's state given default macro property representation parametersMap = self.convertToParameterMap(properties.parameters) return macro_gem.macro_gemProperties( parameter1=parametersMap.get('parameter1') ) def unloadProperties(self, properties: PropertiesType) -> MacroProperties: # convert component's state to default macro property representation return BasicMacroProperties( macroName=self.name, projectName=self.projectName, parameters=[ MacroParameter("parameter1", properties.parameter1) ], ) ``` Similarly the opposite case where this enhanced UX is not available due to some reason, Prophecy needs to be able to render the default macro UI. For this purpose you must define the logic to convert the gem properties back to the default macro properties object which Prophecy understands. ## Preview your gem You can preview the component in the gem builder to see how it looks. You can modify the properties and then save it to preview the generated code which will eventually run on your cluster. Gem builder preview Certain gems may generate SQL code that isn't compatible with a specific fabric provider, rendering the gem unusable and guaranteeing failure if attempted. This issue arises because some dbt macros are designed to support only specific warehouse types. Custom gem logic can be shared with other users within the Team and Organization. Navigate to the gem listing to review Prophecy-defined and User-defined gems. When your gem is ready, publish it so that it is available to use in other models. ## Example code This is an example specification of a gem for an existing deduplicate macro from `dbt utils`. ```sql theme={null} from dataclasses import dataclass from collections import defaultdict from prophecy.cb.sql.MacroBuilderBase import * from prophecy.cb.ui.uispec import * class Deduplicate(MacroSpec): name: str = "deduplicate" projectName: str = "dbt_utils" category: str = "Custom" @dataclass(frozen=True) class DeduplicateProperties(MacroProperties): tableName: str = '' partitionBy: str = '' orderBy: str = '' def dialog(self) -> Dialog: return Dialog("Macro") \ .addElement( ColumnsLayout(gap="1rem", height="100%") .addColumn( Ports(allowInputAddOrDelete=True), "content" ) .addColumn( StackLayout() .addElement( TextBox("Table Name") .bindPlaceholder("Configure table name") .bindProperty("tableName") ) .addElement( TextBox("Deduplicate Columns") .bindPlaceholder("Select a column to deduplicate on") .bindProperty("partitionBy") ) .addElement( TextBox("Rows to keep logic") .bindPlaceholder("Select row on the basis of ordering a particular column") .bindProperty("orderBy") ) ) ) def validate(self, context: SqlContext, component: Component) -> List[Diagnostic]: diagnostics = [] macroProjectMap = self.getMacroMap(context) projectName = self.projectName if self.projectName != "" else context.projectName if projectName not in macroProjectMap: diagnostics.append(Diagnostic( "properties.projectName", f"Project name {self.projectName} doesn't exist. Current Project is ${context.projectName}", SeverityLevelEnum.Error )) else: macroDef: Optional[MacroDefFromSqlSource] = self.getMacro(self.name, projectName, context) if macroDef is None: diagnostics.append(Diagnostic( "properties.macroName", f"Macro {self.name} doesn't exist", SeverityLevelEnum.Error )) else: if component.properties.tableName == '': diagnostics.append( Diagnostic( f"properties.tableName", f"Please define table name", SeverityLevelEnum.Error ) ) if component.properties.partitionBy == '': diagnostics.append( Diagnostic( f"properties.partitionBy", f"Please define partition by column", SeverityLevelEnum.Error ) ) if component.properties.orderBy == '': diagnostics.append( Diagnostic( f"properties.orderBy", f"Please define order by by column", SeverityLevelEnum.Error ) ) return diagnostics def onChange(self, context: SqlContext, oldState: Component, newState: Component) -> Component: return newState def apply(self, props: DeduplicateProperties) -> str: if self.projectName != "": resolved_macro_name = f"{self.projectName}.{self.name}" else: resolved_macro_name = self.name non_empty_param = ",".join([param for param in [props.tableName, props.partitionBy, props.orderBy] if param != '']) return f'{{{{ {resolved_macro_name}({non_empty_param}) }}}}' def loadProperties(self, properties: MacroProperties) -> PropertiesType: parametersMap = self.convertToParameterMap(properties.parameters) return Deduplicate.DeduplicateProperties( tableName=parametersMap.get('relation'), orderBy=parametersMap.get('order_by'), partitionBy=parametersMap.get('partition_by') ) ``` # Stored procedures Source: https://docs.prophecy.ai/data-analysis/development/extensibility/stored-procedure Create and call stored procedures to use in pipelines Stored procedures let you run procedural logic within your pipeline. While most business logic should be implemented using standard Prophecy gems or declarative SQL, stored procedures are useful in specific scenarios where procedural control is required. Use stored procedures when: * Migrating existing logic from systems that use stored procedures. * Running DDL operations, such as creating or cleaning up tables. * Bookkeeping, such as writing execution metadata (run time, parameters, status) to an audit table after a pipeline run finishes. * Iterative operations, such as looping through all tables in a database to extract metadata or perform cleanup tasks. This page describes how to create new stored procedures in Prophecy. For information about calling stored procedures in a pipeline, visit [StoredProcedure](/data-analysis/gems/custom/stored-procedure). ## Prerequisites To use stored procedures, you need: * Prophecy 4.1.2 or later. ## Create stored procedure To create a stored procedure in a project, click **+ Add Entity > Stored Procedure** in the project browser. Then, configure and save the stored procedure. The following table describes the parameters of a stored procedure using the visual view. In the code view, you can write the entire procedure using BigQuery SQL syntax, including the `CREATE PROCEDURE` statement and procedure body. | Parameter | Description | | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Project | The Google Cloud project that contains the dataset where the stored procedure will be created. | | Dataset | The BigQuery dataset where the stored procedure will be created. | | Arguments | Define the parameters passed to the stored procedure.
Each argument includes a name, data type, and mode (`IN`, `OUT`, or `INOUT`). Learn more in [Argument Modes](#argument-modes). | | Code | The body of the stored procedure written in BigQuery SQL. | To learn about and view examples of stored procedures in BigQuery, visit [Work with SQL stored procedures](https://cloud.google.com/bigquery/docs/procedures). ### Argument Modes Stored procedures support the following argument modes: * `IN` mode: Passes a value into the stored procedure. The procedure can read the value but cannot modify it. * `OUT` mode: Returns a value from the stored procedure. The procedure assigns a value to the argument. * `INOUT` mode: Passes a value into the procedure and returns a (possibly updated) value back to the caller. ### Procedure Options Prophecy supports additional options that you can find in the top right corner of the stored procedure configuration. * **Strict Mode**: When `True` (default), the procedure body is checked for errors such as non-existent tables or columns. The procedure fails if the body fails any of these checks. When `False`, the procedure body is checked only for syntax. Learn more about `strictMode` in the [BigQuery API reference documentation](https://cloud.google.com/bigquery/docs/reference/rest/v2/routines). * **Description**: Use this field to document the purpose and behavior of the stored procedure. The description is saved with the procedure in BigQuery. ## Call stored procedure Once you have created a stored procedure, you can call it using the [StoredProcedure gem](/data-analysis/gems/custom/stored-procedure). # Deploy parameter sets Source: https://docs.prophecy.ai/data-analysis/development/parameters/parameter-sets Choose which parameter values are used when pipelines and analyses run A parameter set is a named collection of values for multiple parameters. You can create parameter sets at the project level or the pipeline level. Parameter sets are useful when multiple values must change in sync, such as: * Environment-specific settings (dev, staging, prod) * Regional configurations (US, EU, APAC) * Tenant- or customer-specific values * Currency, locale, or compliance variations Project parameters interface showing default, EU, and APAC parameter sets After you define parameters and group them into parameter sets, you must decide which set to use at execution time. The selected parameter set determines the values Prophecy injects when a pipeline or app runs. ## Set hierarchy When you first open parameter settings, both project parameters and pipeline parameters include an empty default parameter set. The default set serves two purposes: * It defines the list of parameters that exist. * It provides baseline values that other sets can inherit. When you create additional parameter sets: * All parameters from the default set appear automatically. * You provide new values for any parameters you want to override. * Any value you do not override is inherited from the default. For interactive development, the default set is active unless you manually activate another set. ## Create parameter sets Each project and pipeline has a default parameter set. Follow the steps below to create a new parameter set. Open parameter settings from the default set or the pipeline menu. Arrows to the default set and the pipeline
menu * Select either **Project Parameters** or **Pipeline Parameters**. * Add variables to the set by clicking **+ Add Variable**. * Enter values for each variable or choose to inherit from the default set. * Click **Save**. * Click **+ Add Parameter Set**. * Enter a name for the parameter set. * Assign values to the variables or leave them empty to inherit the default. * Toggle **Active** to make this set the active configuration. * Click **Save**. ## Example There are many use cases where you may want to leverage multiple parameter sets. This example demonstrates how to use parameter sets to process sales data for different global regions. When you switch regions, you often need to change more than one piece of information at a time. In this scenario, we use two parameters that must stay in sync: * `currency`: Labels the final report (for example, **USD** or **EUR**). * `fx_rate_table`: Tells the pipeline which database table contains the specific exchange rates for that region. By grouping these into parameter sets, you ensure that when you switch sets, pipelines use both the correct label and the correct exchange rates. ### Define the default values (USA) First, set up your standard operating values in the default set: 1. Open the parameter settings. 2. Click **+ Add Variable** and name it `currency`. 3. Set the value to `USD`. 4. Click **+ Add Variable** again and name it `fx_rate_table`. 5. Set the value to `fx_rates_usd`. 6. Click **Save**. After saving the set of variables, you can reference `currency` and `fx_rate_table` in any pipeline in the project. By default, your pipeline will use these values for interactive runs. ### Create a regional override (Europe) Next, create a set for your European operations: 1. Click **+ Add Parameter Set** and name it `Europe`. 2. In the `Europe` set, change the value of `currency` to `EUR`. 3. Set the value of `fx_rate_table` to `fx_rates_eur`. 4. Toggle **Active** to make this set the active configuration. 5. Click **Save**. By switching the active set to `Europe`, you update your entire project to handle European data without manually changing every pipeline component. ## Select parameter sets at runtime How you select a parameter set depends on how the pipeline or analysis is executed. ### On-demand pipeline runs Interactive pipeline runs use the active parameter set for both project and pipeline parameters. ### On-demand analysis runs You can select a parameter set when [creating an analysis dashboard](/data-analysis/analysis/create-analysis) or modify it afterward in [analysis settings](/data-analysis/analysis/analysis-settings). This determines which pipeline parameter values the analysis uses when it runs. On-demand analysis runs always use values from the active project parameter set. ### Pipeline schedules When you [schedule](/data-analysis/production/scheduling/schedule-setup) a pipeline, you can choose which pipeline parameter set to use for that schedule. This lets you run the same pipeline with different configurations on different schedules. ### Analysis schedules When you [schedule an analysis](/data-analysis/analysis/consume-analysis), you can choose which pipeline parameter set to use for that schedule. This lets you run the same analysis with different configurations on different schedules. ### Deployed pipelines When you [publish](/data-analysis/production/publication) a project, you can choose which project-level parameter set to use. This allows you to publish different configurations to different environments without modifying the project. The selected project parameter set applies to all deployed pipelines in the published project. Pipeline parameter sets selected for schedules also apply to those scheduled runs. If a pipeline parameter and a project parameter share the same name, the pipeline parameter value takes precedence. Deployment activates pipeline schedules so it's all wrapped up in this step. # Parameters in Data Analysis Source: https://docs.prophecy.ai/data-analysis/development/parameters/parameters Reuse values for flexible pipeline execution Parameters let you dynamically control values in pipelines. Instead of hard-coding values into your pipeline, you define them once as parameters and reference them wherever needed. At runtime, you choose which values to use. Use parameters when a value: * Changes between environments (for example, dev vs. prod). * Should differ between pipeline schedules. * Is shared across multiple pipelines. * Will correspond to user-entered fields in [analysis dashboards](/data-analysis/analysis/overview). ## High-level workflow Using parameters in Prophecy follows a simple flow: 1. Define parameters to declare the values your pipeline expects. 2. Reference parameters inside gems in pipelines. 3. Run pipelines using a specific set of parameter values. This page focuses on step 1: defining parameters. ## What are parameters? A parameter is a named placeholder for a value that may change between runs. For example, instead of hard-coding a file path in the pipeline, such as `s3://my-bucket/prod/customers/`, you define a parameter called `customers_path` and assign it the value `s3://my-bucket/prod/customers/`. Inside the pipeline, you reference the parameter rather than the literal value. Parameters are defined in [parameter sets](/data-analysis/development/parameters/parameter-sets), which group related parameters together. Every project and pipeline has a default parameter set, which serves as the baseline definition. You can later create additional parameter setsβ€”for example, one for development and one for productionβ€”without redefining the parameters themselves. ## Types of parameters Prophecy supports two scopes for parameters: * Project parameters * Pipeline parameters The scope determines where a parameter can be used and how it behaves at runtime. | Parameter type | Description | | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Project parameters | Shared across all pipelines in a project. Project parameters support concrete types. Values are interpreted based on their content.

For example, expressions like `concat("first_name", "last_name")` are recognized as valid SQL and evaluated accordingly, rather than treated as literal strings. | | Pipeline parameters | Specific to a single pipeline.

String expressions are evaluated at runtime. If you define a pipeline parameter called `concat("first_name", "last_name")`, the value injected during usage is `first_namelast_name`.

| If a pipeline parameter and a project parameter share the same name, the pipeline parameter overrides the project parameter. ### Parameter storage and editing All parameters are defined in `parameters.py`. If a parameter can be represented in `dbt.yaml`, it is also written there. Otherwise, it remains only in `parameters.py`. This enables support for richer data types that are not compatible with YAML. Parameters can be edited through: * The visual editor * `parameters.py` * `dbt.yaml` (when supported) ### Supported data types Prophecy supports typed parameters at both the project and pipeline level. | Data type | Description | | -------------- | ------------------------------------------------------------------------------------------------------------ | | Array | A list of values of the same type, added one by one in a multi-value input field. Stored in `parameters.py`. | | Boolean | A true or false value, selected from a dropdown. | | Connection | A connection in the current [fabric](/data-analysis/environment/fabrics/prophecy-fabrics). | | Date | A calendar date in `dd-mm-yyyy` format, chosen using a date picker. Stored in `parameters.py`. | | Double | A 64-bit floating-point number entered in a numeric field. | | Float | A 32-bit floating-point number entered in a numeric field. | | Int | A 32-bit integer entered in a numeric field. | | Long | A 64-bit integer entered in a numeric field. | | SQL Expression | A SQL expression, either configured through dropdown menus or entered as custom code. | | String | A plain text value, entered via a single-line text input. | ## Define parameters You intiially define parameters by adding variables to the default parameter set. You must define values for default parameters. To define project or pipeline parameters: 1. Open the parameter settings. For pipeline parameters, you must open the relevant pipeline first.Open parameter settings 2. Select either the **Project Parameters** or **Pipeline Parameters** tab. 3. Click **+ Add Variable**. 4. Enter a name and a value for the variable. 5. Click **Save**. ### Usage (precedence and selection) At runtime, you can use both project and pipeline parameters. When multiple values are defined, parameters are resolved in the following order (lowest to highest priority): 1. Default project parameter set 2. Custom project parameter set 3. Default pipeline parameter set 4. Custom pipeline parameter set ### Connection parameters Connection parameters reference fabric connections rather than storing scalar values such as strings or numbers. You can use connection parameters to make pipelines portable across environments while preserving connection-type safety. Connection parameters are: * Scoped to the current fabric. * Restricted to compatible connection types. * Validated against available fabric connections. * Revalidated when switching fabrics. See the [connection parameters example in Use Parameters](/data-analysis/development/parameters/usage#connection-type) for an implementation example. ### Migration behavior Project parameters have been enhanced as as version `4.2.7`. If you created project parameters in earlier versions, Prophecy automatically migrates them to the new system. During migration: * Values in double quotes are treated as strings * Numeric values are inferred as numeric types * Values are parsed using a SQL parser: * If valid SQL β†’ stored as SQL expression * Otherwise β†’ treated as string ## What's next? Now that you understand how parameters are defined, continue with: * [Using these parameters](/data-analysis/development/parameters/usage) in your pipelines and analyses. * [Creating additional parameter sets](/data-analysis/development/parameters/parameter-sets) at the project or pipeline level. # Use parameters Source: https://docs.prophecy.ai/data-analysis/development/parameters/usage Reference parameters in pipelines and analysis dashboards Parameters let you define reusable variables that are injected into your pipeline at runtime. Instead of hard-coding values such as dates or file paths, you can reference parameters. This page describes how to operationalize existing parameters in your pipeline for [analysis dashboards](/data-analysis/analysis/overview) or [deployment](/data-analysis/production/publication). To learn how to create and define new parameters, see the [parameters overview](/data-analysis/development/parameters/parameters). ## How do you use parameters in your pipeline? Once you create parameters, they are available as *configuration variables* in gems. You can also reference parameters using Jinja syntax in code expressions. To access parameters in a gem: 1. Open any [gem](/data-analysis/gems/) that uses visual or code expressions, such as a [Filter](/data-analysis/gems/prepare/filter) or [Reformat](/data-analysis/gems/prepare/reformat). 2. In Visual mode, select **Configuration Variables** from the visual expression builder. You'll see a list of all existing parameters in your project. 3. In Code mode, use Jinja syntax instead. Use the following syntax to reference a parameter: `{{ var('parameter_name') }}`. You can also use Jinja syntax in the visual expression builder by including a custom code expression. For more information on Jinja variables, see [dbt's documentation on Jinja](https://docs.getdbt.com/reference/dbt-jinja-functions/var). Prophecy determines the parameter's value based on the [active parameter set](/data-analysis/development/parameters/parameter-sets). By default, Prophecy uses the values in the default parameter set. ## Examples ### Array type This example uses a dataset with a column called `region`. You can use an **Array** parameter called `region_list` to filter rows that match one of several regions. #### Create the Array parameter 1. Open the pipeline parameter settings for the relevant pipeline. 2. Click **+ Add Variable**. 3. Name the parameter `region_list`. 4. Select the **Type** and choose **Array**. 5. Select **String** for **Array** type. 6. Click **+** to add items to the array. 7. Click **Value** and enter `US-East` (or another region code). 8. Click **Done**. 9. Repeat steps 6-8 to add `US-West` and `Europe` to the **Array** parameter. 10. Click **Save**. #### Use the parameter in a filter Next, you'll filter your dataset to only include rows where the `region` column matches a value in `region_list`. 1. Add a **Filter** gem to your pipeline. 2. Remove the default `true` expression. 3. Click **Select expression**. 4. Select **Function > Array > array\_contains**. 5. Choose **value > Configuration Variable**. 6. Select `region_list`. 7. Click **+** to add an argument for `array_contains` and choose `Region`. 8. Click **Save**. 9. Add a Target table gem called `sales_transactions_by_region` and connect it to the Filter gem. 10. Click **Save**. #### Adjust region from a dashboard You can now select the parameter in an [analysis dashboard](/data-analysis/analysis/overview) so that end users can select the regions they want to see in a report. 1. Create an [analysis](/data-analysis/analysis/create-analysis) for the `regional_sales` pipeline. 2. Add a title for the analysis. 3. Select **Interactive > Checkbox Group**. 4. Select `region_list` for **Configuration field**. 5. Add a region in **Default value**, such as `US-East`. 6. Add options such as `US-East`, `US-West`, `Europe`, `Mexico`, `Brazil`, `LAC`, and `Andean`. 7. Open the **Data Integration** dropdown and select **Data Preview**. 8. In the **Inspect** tab, choose `sales_transactions_by_region`. 9. Select columns to display. When the analysis runs, users can check boxes to select their desired regions. For example, a sales team in Latin America might select `Mexico`, `LAC`,`Brazil`, and `Andean` to view their focus regions. ### Date type This example uses a dataset with timestamped sales data. You can use two **Date** parameters, `start_date` and `end_date` to configure a snapshot of sales data by a time period such as week or month. #### Create the Date parameters 1. Open the pipeline parameter settings for the relevant pipeline. 2. Click **+ Add Parameter**. 3. Name the parameter `start_date`. 4. Select **Type** and choose **Date**. 5. Click **Select expression > Value**. 6. Enter `09/01/2025` (or another default start date) and click **Done**. 7. Click **Save**. 8. Repeat the steps above to create an `end_date` parameter with a default value of `09/07/2025`. #### Use the parameters in a filter 1. Add a **Filter** gem to your pipeline. 2. Remove the default `true` expression. 3. Click **Select expression**. 4. Select **Column** and select `sales_date` (or your dataset's date column). 5. Choose the **between** operator. 6. For both `start_date` and `end_date`, click **Select expression > Configuration Variable** and select corresponding parameters. 7. Add a Target table gem called `snapshot_by_date` and connect it to the Filter gem. 8. Click **Save**. #### Adjust date from a dashboard 1. Create an [analysis](/data-analysis/analysis/create-analysis) for the `sales_snapshot` pipeline. 2. Add a title. 3. Select **Interactive > Date Field**. 4. For **Configuration field**, choose `start_date`. 5. Add another **Date Field** and select `end_date`. 6. Open the **Data Integration** dropdown and select **Data Preview**. 7. In the **Inspect** tab, choose `snapshot_by_date`. 8. Select columns to display. When the analysis runs, users can select their own values for `start_date` and `end_date`. ### String type This example uses a dataset with a column called `customer_category` with values such as `Premium`, `Basic`, and `Standard`. You can use a **String** parameter called `customer_type` to filter rows for a specific group of customers. #### Create the String parameter 1. Open the pipeline parameter settings for the relevant pipeline. 2. Click **+ Add Parameter**. 3. Name the parameter `customer_type`. 4. Select the **Type** and choose **String**. 5. Click **Select expression > Value**. 6. Enter `Premium` and click **Done**. 7. Click **Save**. #### Use the parameter in a filter Next, you'll filter your dataset based on the `customer_type` parameter. 1. Add a **Filter** gem. 2. Remove the default `true` expression. 3. Click **Select expression > Column** and select `customer_category`. 4. Choose the **Equals ( = )** operator. 5. Click **Select expression > Configuration Variable**. 6. Select `customer_type`. 7. Add a Target table gem called `filtered_customers` and connect it to the Reformat gem. 8. Click **Save**. #### Adjust customer type from a dashboard 1. Create an [analysis](/data-analysis/analysis/create-analysis) for the `customer_segment` pipeline. 2. Add a title for the analysis. 3. Select **Interactive > Dropdown**. 4. Give the dropdown a label. 5. Select `customer_type` for **Configuration field**. 6. Open the **Data Integration** dropdown and select **Data Preview**. 7. In the **Inspect** tab, choose `filtered_customers`. 8. Select the columns to display. When the analysis runs, users can switch the `customer_type` parameter from `Premium` to `Standard` (or another category) to explore different customer groups. ### Boolean type This example uses a dataset of customer reviews, in which reviews older than 5 years are designated as `archived`, using a column called `archived_reviews` with Boolean values. You can use a Boolean parameter to create an analysis dashboard that lets users choose whether to include archived reviews. #### Create the Boolean parameter 1. Open the pipeline parameter settings for the relevant pipeline. 2. Click **+ Add Parameter**. 3. Name the parameter `include_archived`. 4. Select the **Type** and choose **Boolean**. 5. Click **Select expression > Value**. 6. Click **False** and click **Done**. 7. Click **Save**. #### Use the parameter in a filter Next, you'll create a Filter gem that uses the `include_archived` parameter in an expression. 1. Create and open the **Filter** gem. 2. Remove the default `true` expression. 3. Click **Select expression > Column** and select `archived`. 4. In the **Select operator** dropdown, select **equals**. 5. In the **Select expression** dropdown of the Filter condition, select **Configuration variable** and select `include_archived`. 6. Add a Target table gem called `prod_filtered_archived` and connect it to the Filter gem. 7. Click **Save**. The output of this gem will only include rows where `include_archived` is false. In the steps below, you'll create an analysis dashboard that lets users change `include_archived` to true. #### Adjust reviews from a dashboard 1. Create an [analysis](/data-analysis/analysis/create-analysis) for the `reviews` pipeline. 2. Add a **Title** for the analysis. 3. Add a **Toggle** that uses `include_archived` as a **Configuration** field, with a label reading `Include archived reviews?`. 4. Open the **Data Integration** dropdown and select **Data Preview**. 5. In the **Inspect** tab, choose `prod_filtered_archived` for **Data table**. 6. Select columns to display. When the analysis runs, users can toggle `Include archived reviews?` to include archived reviews in results. ### Double type This example uses a dataset that includes a column called `discount_rate` that applies a discount for customers in certain cases. You can use a Double parameter inside an analysis dashboard that lets users adjust this rate. #### Create the Double parameter 1. Open the pipeline parameter settings for the relevant pipeline. 2. Click **+ Add Parameter**. 3. Name the parameter `discount_rate`. 4. Select the **Type** and choose Double. 5. Click **Select expression > Value**. 6. Enter `.15` and click **Done**. 7. Click **Save**. #### Use the parameter in a reformat Next, you'll create a Reformat gem that uses the `discount_rate` parameter in an expression that uses Jinja syntax. 1. Add a [Reformat gem](/data-analysis/gems/prepare/reformat). 2. Under **Target Column**, add `price`, `product`, and `quantity`. 3. Under **Target Column**, add a new column called `discounted_price`. 4. Click **Select expression > Custom code** and enter `price * (1 - {{ var('discount_rate') }})`. 5. Add a Target table gem called `products_discounted` and connect it to the Reformat gem. 6. Click **Save**. #### Adjust discount rate from a dashboard 1. Create an [analysis](/data-analysis/analysis/create-analysis) for the `products_with_reviews` pipeline. 2. Add a title for the analysis. 3. Select **Interactive > Number Input**. 4. Select `discount_rate` for **Configuration field**. 5. Give the field a label. 6. Open the **Data Integration** dropdown and select **Data Preview**. 7. In the Inspect tab, select `products_discounted` for **Data table**. 8. Select columns to display. When the analysis runs, users can enter their own rate for `discount_rate`. ### Long type This example uses a dataset for a telecom company that includes aggregated usage data by month. You can use a `Long` parameter to set a monthly data cap in MB and flag or filter subscribers who exceed it. #### Create the Long parameter 1. Open the pipeline parameter settings for the relevant pipeline. 2. Click **+ Add Parameter**. 3. Name the parameter `usage_cap_mb`. 4. Select the **Type** and choose **Long**. 5. Click **Select expression > Value**. 6. Enter `50000` and click **Done**. 7. Click **Save**. #### Use the parameter in a filter 1. Add a **Filter** gem. 2. Remove the default `true` expression. 3. Select **Column > total\_usage\_mb**. 4. Choose **Greater than ( > )**. 5. Click **Select expression > Configuration Variable** and select `usage_cap_mb`. 6. Add a **Table** gem called `usage_over_cap` and connect it to the **Filter** gem. 7. Click **Save**. #### Adjust usage cap from a dashboard 1. Create an [analysis](/data-analysis/analysis/create-analysis) for the `usage_cap_monitor` pipeline. 2. Add a title for the analysis. 3. Select **Interactive > Number Input**. 4. Select `usage_cap_mb` for **Configuration field** and label it **Monthly Cap (MB)**. 5. Open **Data Integration > Data Preview**. 6. In the **Inspect** tab, choose `usage_over_cap` for **Data table**. 7. Select columns to display (e.g., `subscriber_id`, `total_usage_mb`, `billing_period`). When the analysis runs, users can raise or lower the cap by changing `usage_cap_mb` to see which subscribers are affected. ### Float type This example uses dataset of sensor data with a column called `sensor_temp`. You can use a **Float** parameter called `temperature_threshold` to filter out rows below a certain temperature. #### Create the Float parameter 1. Open the pipeline parameter settings for the relevant pipeline. 2. Click **+ Add Parameter**. 3. Name the parameter `temperature_threshold`. 4. Select the **Type** and choose **Float**. 5. Click **Select expression > Value**. 6. Enter `72.1` and click **Done**. 7. Click **Save**. #### Use the parameter in a filter Next, you'll use the `temperature_threshold` parameter to filter your data. 1. Add a **Filter** gem. 2. Remove the default `true` expression. 3. Select **Column > sensor\_temp**. 4. Choose the **Greater than ( > )** operator. 5. Click **Select expression > Configuration Variable**. 6. Select `temperature_threshold`. 7. Add a Table gem called `filtered_temperature` and connect it to the Filter gem. 8. Click **Save**. #### Adjust temperature threshold from a dashboard 1. Create an analysis for the `temperature_monitor` pipeline. 2. Add a title for the analysis. 3. Select **Interactive > Number Input**. 4. Select `temperature_threshold` for **Configuration field**. 5. Give the field a label, such as **Temperature Threshold**. 6. Open the **Data Integration** dropdown and select **Data Preview**. 7. In the **Inspect** tab, choose `filtered_temperature`. 8. Select columns to display. When the analysis runs, users can adjust `temperature_threshold` to make filtering more or less sensitive. ### Connection type This example uses a pipeline that reads data from Amazon S3. You can use a **Connection** parameter to make the pipeline reusable across fabrics and environments. #### Create the Connection parameter 1. Open the pipeline parameter settings for the relevant pipeline. 2. Click **+ Add Variable**. 3. Name the parameter `sales_s3_connection`. 4. Select the **Type** and choose **Connection**. 5. Select a connection from the current fabric, such as `s3_sales_dev`. 6. Click **Save**. The selected connection becomes the default value for the parameter. #### Use the parameter in a Source gem Next, you'll configure a Source gem to use the connection parameter instead of a fixed connection. 1. Add an **S3 Source** gem to your pipeline. 2. Open the Source gem settings. 3. In the **Connection** field, select **Configuration Variable**. 4. Select `sales_s3_connection`. 5. Configure the remaining source settings, such as bucket and file path. 6. Click **Save**. Only connection parameters matching the Source gem type are selectable. For example, an S3 Source gem only allows S3 connection parameters. #### Switch fabrics If you switch to another fabric, Prophecy validates that the configured connection still exists and matches the expected type. Diagnostics appear in the Source gem if: * The selected connection does not exist in the current fabric. * A connection with the same name exists, but the connection type does not match the Source gem type. For example, if `sales_s3_connection` references an S3 connection in one fabric, but the same connection name refers to a Snowflake connection in another fabric, the Source gem displays a diagnostic error. ## Best practices To make the most out of parameters, we suggest you: * Use meaningful parameter names that indicate their purpose. * Validate inputs to prevent unexpected errors during execution. * Keep sensitive values such as API keys in [secrets](/data-analysis/environment/secrets/secrets) rather than passing them as plain parameters. # Pipelines for Data Analysis Source: https://docs.prophecy.ai/data-analysis/development/pipelines/data-analysis-pipelines Data preparation workflows for transforming warehouse data Pipelines are visual workflows that define how data is transformed inside a project. Pipelines consist of a sequence of gems that collect, transform, and move data from sources to storage or analytics systems. For example, a pipeline might join orders and customer tables, standardize date formats, calculate revenue metrics, and materialize a curated reporting table. Instead of writing isolated SQL queries, you define a transformation graph that: * Encodes business logic in a visual workflow. * Can be iteratively refined. * Produces consistent, reusable outputs. * Can be versioned. Because pipelines compile to SQL, all execution happens in your warehouse under your existing credentials and access controls. In Prophecy's **Data Prep and Analysis** model: * **Pipelines** handle data preparation. * [**Analyses**](/data-analysis/analysis/overview) handle exploration, visualization, and interpretation. ## Working with pipelines A pipeline contains a connected sequence of gems that you can edit in [Prophecy's Studio](/data-analysis/development/studio/studio). See [Data analysis gems](/data-analysis/gems/gems) for more information on gems. Pipelines appear in the project browser alongside other project artifacts. Pipeline example ## Build a pipeline ### Build a pipeline with the Transform Agent The [Transform Agent](/data-analysis/ai/agent/transform) can generate or modify pipelines using natural language. The Agent accelerates development, but all transformations remain visible and editable. When using the Agent, you: * Describe the transformation you want. * Review generated gems and connections. * Inspect compiled SQL. * Validate results before finalizing changes. ### Build a visually in the pipeline canvas You can also create a pipeline from the project landing page or the [Project Browser](/data-analysis/development/studio/studio#project-browser). When creating a pipeline, you provide: * **Pipeline name** β€” A unique name within the project. * **Directory path** β€” The location where the pipeline is stored in the project. ## Open pipeline To open a pipeline, click its name in the Project Browser or on the project landing page. ## Modify pipeline Once you have created a pipeline, you can: * Drag gems onto the canvas and connect gems to define data flow. * Configure gems to produce desired output. * Preview results using the [Data Explorer](/data-analysis/development/runs/data-explorer/data-explorer). * Inspect SQL in code view. * Run the pipeline. ## Rename pipeline To rename a pipeline: 1. Hover over the pipeline in the Project Browser. 2. Click the **Rename** icon. 3. Enter a new pipeline name and click **Rename**. Rename pipeline ## Duplicate pipeline You can duplicate entire pipelines. To duplicate a pipeline: 1. Open the Pipeline in Studio. 2. Click **...** menu in the upper-right corner of the [pipeline canvas](/data-analysis/development/studio/studio#pipeline-canvas). 3. Click **Duplicate** in the dialog box. The duplicated pipeline opens in the pipeline canvas. By default, the pipeline is named `_copy`. Duplicate pipeline menu ## Mark pipeline as draft You can mark a pipeline as `draft`. This lets you keep working on it, while excluding changes from the project's next published version. For more on how this affects publication, see [Draft pipelines](/data-analysis/production/publication/publication-concepts#draft-pipelines). To mark a pipeline as draft: 1. Open the pipeline in Studio. 2. Click the **...** menu in the upper-right corner of the pipeline canvas. 3. Click **Keep as draft**. 4. Confirm the change. While a pipeline is a draft: * Its in-progress changes aren't included in the next published version. If it's already been published before, it ships unchanged in the new version; if it hasn't, it's left out of the version entirely. * Existing schedules on the pipeline keep running β€” publishing doesn't take it down or interrupt it. * You can't rename or delete the pipeline. To remove draft status: 1. Open the pipeline in Studio. 2. Click the **...** menu in the upper-right corner of the pipeline canvas. 3. Click **Remove draft**. 4. Confirm the change. Removing draft status doesn't publish the pipeline immediately β€” it's included the next time the project is published. Draft pipelines are available only for projects that use Simple Version Control. ## Execution and validation Pipelines move through development, production, and reporting stages. You can [run them interactively in the canvas](/data-analysis/development/runs/execution), schedule automated runs, or trigger execution through analyses. When you run a pipeline: * Each transformation is translated into warehouse-native SQL that runs directly in your connected warehouse. * Execution respects your existing permissions. * Output datasets are updated based on the defined transformations. You can re-run pipelines to validate changes or refresh prepared datasets before analysis. You can also run individual gems to validate intermediate outputs. ## Human in the loop Whether you build pipelines manually or generate them with the Transform Agent, pipelines remain under your control. You can: * Inspect every transformation. * Modify generated logic. * Validate schema changes. * Re-run transformations as needed. ## Export code You can export the SQL for a single pipeline or all the pipelines in a project. See [Export compiled code](/data-analysis/development/projects/export-code) for more details. # View pipeline Python Source: https://docs.prophecy.ai/data-analysis/development/pipelines/view-pipeline-py Understand how the pipeline.py file defines pipeline structure and execution order. The `pipeline.py` file defines the structure of a pipeline as code. It acts as the code-based representation of the visual pipeline graph, allowing you to understand how steps are connected and executed. Previous versions of this file used task-based structure; Prophecy now uses a graph-based model for this file. Instead of organizing logic by tasks, the pipeline is now defined as a set of processes connected through dependencies, which determine execution order. ## Overview The `pipeline.py` file serves as the source of truth for generating the pipeline graph. It defines how pipeline steps relate to each other. There is not always a one-to-one mapping between gems and nodes. Some gems may be grouped into an execution unit. In this model: * Nodes (vertices) represent steps in the pipeline, such as data sources, transformations, models, or outputs. * Edges represent dependencies between steps, indicating execution flow. ## How the pipeline is defined The file defines pipelines using a declarative graph structure. * Nodes are represented by `Process` objects. * Edges (connections) are created using the `>>` operator. * The graph structure is captured using context management. For example: ``` source >> transform >> sink ``` defines the following execution flow: ``` source β†’ transform β†’ sink ``` ### Example The following snipped represents a pipeline that runs a transformation (`sales_by_region`) and then sends the results via email. The connection `transform >> email` shows that the email step depends on the transformation output. See [classes](#classes) for an explanation of classes. ```Python theme={null} with Pipeline(args) as pipeline: transform = Process( name="sales_by_region", properties=ModelTransform(modelName="sales_by_region") ) email = Process( name="send_report", properties=Email(subject="Sales ΰ€°ΰ€Ώΰ€ͺΰ₯‹ΰ€°ΰ₯ΰ€Ÿ", to="team@example.com") ) transform >> email ``` ## How to access the file You can view the `pipeline.py` file in the [Project Browser](/data-analysis/development/studio/studio#project-browser/) while in [Code view](/data-analysis/development/studio/studio#visual-code-toggle). 1. Go to **Project**. 2. Select **Pipelines**. 3. Open the `.py` file listed under **Pipelines**. ## How to use this file You can use the `pipeline.py` file to: * Understand pipeline structure. * Determine execution order based on dependencies. * Inspect how processes are connected in the graph. This file is informational and reflects how the pipeline is defined internally. ## How it differs from the previous model Previously, pipelines were organized by tasks, and execution order could be inferred from the task structure. In the current model: * The structure is organized by process instead of task. * Execution order is determined by the dependency graph between processes. * The underlying execution logic has not changed. * Only the representation has changed from task-based to graph-based. ## Relationship to the visual pipeline The `pipeline.py` file is the code counterpart of the visual pipeline graph. * It provides the structural definition used to render the graph. * You can use it to understand how different steps are connected. * It reflects execution flow through explicit dependencies. * Gems generally map to processes. * In some cases, multiple gems may be grouped into a single execution unit instead of a one-to-one mapping. ## CI/CD considerations This file is not intended for CI/CD usage. ## Editing the file You can edit the `pipeline.py` file directly, but we recommend using the Agent to modify it. ## Classes | Class | What it Represents | What to Look For | | ---------------------- | -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **`Pipeline`** | The overall pipeline definition | Wraps the entire pipeline. Everything inside defines the pipeline structure. Think of this as β€œthe full workflow” | | **`PipelineArgs`** | Basic metadata about the pipeline | `label`: pipeline name

`version`: pipeline version

Optional settings (such as `layout`) | | **`Process`** | A step in the pipeline (may combine multiple gems) | Each process is one operation (such as `Transform`, `Visualize`, or `Email`).

`name` identifies the step- `properties` defines what the step does. | | **Connections (`>>`)** | The flow of data between steps | `A >> B` means β€œA feeds into B.”

Defines order and dependencies.

Chains show the path data follows. | # View SQL lineage Source: https://docs.prophecy.ai/data-analysis/development/pipelines/view-sql-lineage Track where output columns come from and how they were transformed in Prophecy. SQL lineage helps you understand where data comes from in a pipeline. It traces which source table columns produced each output column and shows how those columns were transformed along the way. Use SQL lineage when you need to inspect how a selected column was created, whether it passed through gems unchanged, or where it was modified in the pipeline. ## View column lineage in a pipeline To view column lineage: 1. Select **Lineage** from the pipeline menu. A lineage panel opens at the bottom of the pipeline editor. 2. Select a dataset and column to see how the column was produced within the pipeline. Prophecy highlights the gem where the selected column is created and displays the transformation logic used to generate it. For example, selecting `AVG_SPECIALIST_VISITS` highlights the Aggregate gem and shows the expression `ROUND(AVG(specialist_visits_expected), 2)` used to create the column. view sql lineage ## Limitations Prophecy cannot compute lineage for some Custom gems, such as the [Rest API gem](/data-analysis/gems/custom/rest-api) or the [Script gem](/data-analysis/gems/custom/script). Because these gems can contain arbitrary code, lineage information may not be available for columns processed within them. # Create a project Source: https://docs.prophecy.ai/data-analysis/development/projects/create-project How to create a new project in Prophecy To get started with Prophecy, you need to create a new project. This is where you will build your pipelines. ## Your first project You'll create your first project using Prophecy's default [project creation template](/data-analysis/administration/management/teams/settings/project-creation-template). You won't configure much β€” just give your project a name and select a team. 1. Click on the **Create Entity** button in the left navigation bar. 2. Hover over the **Project** tile and select **Create**. 3. Give your project a name. 4. Under **Team**, select a team. We recommending using your personal team (matching your email address) for your first project, so that only you will have access to the project. 5. Under **Select Template**, choose **Prophecy for Analysts**. 6. Click **Complete**. Prophecy opens the Studio. For a new project with v4 AI enabled, Agent chat opens maximized by default. You can minimize the chat and create entities from the project landing page or Project Browser. If you do not have any existing fabrics, you'll need to create one before running pipelines. ## Common questions Yes, you can host the project in your own Git repository. To do so, create a new project using the following steps: * Click on the **Create Entity** button in the left navigation bar. * Hover over the **Project** tile and select **Create**. * Give your project a name. * Under **Team**, select your personal team. * Under **Select Template**, choose **Custom**. **This differs from the above steps.** * Click **Continue** to open the **Git Repository** tab. * Choose existing Git credentials or connect new Git credentials. * Under **Repository**, specify the repository to store the project code. The repository must be empty. * Under **Default Branch**, select the branch that will be the production branch. * Optional: Under **Path**, specify a path in the repository to store the project code. The directory must be empty. * Under **Development Branch**, select the branch that will contain all working changes. * Click **Continue** to complete the project creation. In the Free and Professional Editions, Prophecy only supports GitHub as an external Git provider. Prophecy for Analysts is a project creation template designed specifically for data analysts. When you select this template, Prophecy automatically configures your project with Prophecy-managed Git in Simple mode and initializes it for a Databricks SQL warehouse. This template provides the most streamlined experience for users who primarily work with SQL and prefer visual interfaces over complex Git workflows. # Export compiled code Source: https://docs.prophecy.ai/data-analysis/development/projects/export-code Export compiled SQL for an entire project or a single pipeline as a ZIP file. Export compiled code lets you download the SQL generated from your pipelines. This is useful when you want to review the SQL produced by Prophecy, run it outside Prophecy, or share it with other teams working directly in the warehouse. You start the export from the project menu in the Studio header. During export, you choose whether to download code for the entire project or for a single pipeline, and Prophecy packages the compiled SQL into a ZIP file. ## Export compiled code 1. Open the project settings menu beside your project’s name in the Studio header. 2. Select **Export Compiled Code** from the menu. 3. In the dialog box, choose whether to download code for the entire project or for a single pipeline. 4. Click **Download Zip File**. The ZIP file contains compiled SQL for the selected project or pipeline. ## Exported file structure The downloaded ZIP file contains a folder for the project. Inside the project folder, each pipeline appears in its own subfolder. Each pipeline folder contains one or more SQL files generated from the pipeline. Example structure: ``` my_project/ test_pipeline/ test_pipeline__customer_culture_join.sql test_pipeline__demographics_culture.sql test_pipeline__popular_song_by_country.sql ``` SQL files use the naming pattern: ``` __.sql ``` The first part of the file name is the pipeline name. The second part corresponds to the gem used to generate the compiled SQL. When a pipeline branches, multiple SQL files can be generated. Each file represents compiled SQL for a specific path in the pipeline graph. For example: * `test_pipeline__demographics_culture.sql` corresponds to the `demographics_culture` aggregation gem. * `test_pipeline__popular_song_by_country.sql` corresponds to the `popular_song_by_country` gem. * `test_pipeline__customer_culture_join.sql` corresponds to the `customer_culture_join` join gem that feeds downstream branches. # Project languages Source: https://docs.prophecy.ai/data-analysis/development/projects/project-languages Supported languages for projects Prophecy supports two types of projects for data analysis. Hop to each section to learn more about each project type. The project type corresponds to the language Prophecy compiles your visual project components into. * [SQL](#sql) * [PySpark](#simplified-pyspark) ## SQL Prophecy privileges SQL as the primary language for code generation for data analysis. Prophecy supports a variety of SQL warehouses to execute pipeline-generated SQL queries. For most use cases, use SQL as project language. ## Simplified PySpark Private Preview Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Simplified PySpark is a project type that abstracts PySpark into a simple, data analysis interface. Use PySpark when your organization uses PySpark as the production codebase language. This makes the collaboration between data analysts and data engineers easier, as they can both use the same language. When you use PySpark, Prophecy generates Python files instead of SQL files. The visual interface remains unchangedβ€”you continue working with SQL expressions and the same gem configurations. ### Requirements To use the Simplified PySpark project type, you need: * A Databricks Spark cluster defined in your fabric. Simplified PySpark does not work with Databricks serverless or any other Spark provider, such as Amazon EMR. * A Databricks SQL warehouse configured in your fabric. This is if you want to switch your project to SQL later. * An init script in Databricks that installs required libraries on your cluster. The library installs functionality that is usually executed by Prophecy Automate. ``` #!/bin/bash set -eu echo 'GLIBC_TUNABLES=glibc.rtld.optional_static_tls=16384' > /etc/environment WHEEL="/Workspace/Shared/prophecy_automate/artifacts/prophecy_automate-2.0.0-py3-none-manylinux_2_17_x86_64.whl" python3 -m pip install --no-cache-dir "$WHEEL" ``` Once you create a Simplified PySpark project, you need to manually start the cluster **in Databricks** to run pipelines. There is currently no way to start a cluster or select a different cluster in Simplified PySpark projects. ### How to create a PySpark project The steps to create a PySpark project are similar to creating a SQL project. The only difference is the project type. 1. Click on the **Create Entity** button in the left navigation bar. 2. Hover over the **Project** tile and select **Create**. 3. Give your project a name. 4. Under **Team**, select your personal team. (It will match your individual user email.) 5. Under **Select Template**, choose **Custom**. 6. For the **Project Type**, choose **Spark/Python (PySpark) > Simplified**. 7. Click **Continue**. 8. Under **Connect Git Account**, connect to an external Git provider or select Prophecy-managed Git. 9. Click **Continue**. To save your project creation configuration, create a [project creation template](/data-analysis/administration/management/teams/settings/project-creation-template) that you can reuse for future projects. ## Comparison matrix While we're working to achieve full parity between languages, there are some differences between SQL and PySpark. The following table compares the feature availability between the two languages. | Feature | Comparison | | ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Gems | Most transformation gems work identically in SQL and PySpark modes. Some gems gain additional capabilities in PySpark. For example, SQL can only write to one output at a time, while PySpark can write to multiple outputs. Some gems may not have PySpark implementations. If you switch to PySpark and use a gem without a Python implementation, Prophecy displays an error indicating that the gem doesn't support PySpark. | | Orchestration | PySpark pipelines can still be orchestrated using Prophecy Automate scheduling. However, you can also schedule using Databricks jobs. If you switch to SQL from PySpark, your project will not be compatible with Databricks jobs. | | Tests | SQL projects support data tests. PySpark projects support data tests and unit tests. | ### Switch between SQL and PySpark To switch between SQL and PySpark, edit the backend language in the project's [development settings](/data-analysis/development/studio/development-settings). Note that not every feature or configuration translates perfectly between languages. Always save or commit your work before switching so you can easily revert if something doesn't translate. # Area chart Source: https://docs.prophecy.ai/data-analysis/development/runs/data-explorer/charts/area-chart Learn what parameters you need to configure an area chart An area chart is a type of line chart that features filled areas under the lines to represent quantitative data over time or across categories. AreaChart You can configure the following parameters for the chart: | Parameter | Description | | -------------------- | -------------------------------------------------------------- | | X-axis column | Column used for the X-axis values, typically time or sequence. | | Y-axis column | Column with aggregated values used for the Y-axis. | | Min. Value | Minimum value displayed on the Y-axis. | | Max. Value | Maximum value displayed on the Y-axis. | | Tick Interval | Interval between tick marks on the Y-axis. | | Enable Gradient Fill | Whether to apply a gradient fill under the area curve. | | Stack Data Series | Whether to stack multiple data series on top of one another. | | Display Legends | Whether to display the legend on the chart. | | Enable Tooltips | Whether to display tooltips on hover. | | Show Grid Lines | Whether to display grid lines on the chart. | # Bar chart Source: https://docs.prophecy.ai/data-analysis/development/runs/data-explorer/charts/bar-chart Learn what parameters you need to configure a bar chart A bar chart groups data by categories, using rectangular bars with heights or lengths proportional to the values they represent. BarChart You can configure the following parameters for the chart: | Parameter | Description | | --------------- | ------------------------------------------- | | X-axis column | Column used for the X-axis categories. | | Y-axis column | Column used for the Y-axis values. | | Min. Value | Minimum value displayed on the Y-axis. | | Max. Value | Maximum value displayed on the Y-axis. | | Tick Interval | Interval between tick marks on the Y-axis. | | Display Legends | Whether to display the legend on the chart. | | Enable Tooltips | Whether to display tooltips on hover. | | Show Grid Lines | Whether to display grid lines on the chart. | # Candlestick chart Source: https://docs.prophecy.ai/data-analysis/development/runs/data-explorer/charts/candlestick-chart Learn what parameters you need to configure a candlestick chart A candlestick chart is a financial chart that displays how prices change for an asset over time, such as stocks and currency. Each candlestick represents the opening, closing, high, and low prices for that time period. You can configure the following parameters for the chart: | Parameter | Description | | ------------------------------- | ----------------------------------------------------------------- | | X-axis column | Column used for the X-axis values, typically a timestamp or date. | | Choose column for open price | Column containing the opening price. | | Choose column for close price | Column containing the closing price. | | Choose column for lowest price | Column containing the lowest price. | | Choose column for highest price | Column containing the highest price. | | Display Legends | Whether to display the legend on the chart. | | Enable Tooltips | Whether to display tooltips on hover. | # Charts Source: https://docs.prophecy.ai/data-analysis/development/runs/data-explorer/charts/charts View charts of data samples between gems Charts let you visualize your data at different stages of the pipeline. To view a chart, open a data sample and navigate to the **Visualization** tab of the [Data Explorer](/data-analysis/development/runs/data-explorer/data-explorer). Charts are based on the data currently loaded in the sample. VisualizationView ## Chart types You can create the following charts in the **Visualization** tab. | Chart Type | Description | | ------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | | [Bar](/data-analysis/development/runs/data-explorer/charts/bar-chart) | Group your data by categories with rectangular bars with heights or lengths proportional to the values that they represent. | | [Line](/data-analysis/development/runs/data-explorer/charts/line-chart) | Visually represent your data over time or along a continuous range. | | [Area](/data-analysis/development/runs/data-explorer/charts/area-chart) | Line chart with filled areas under the lines to represent quantitative data over time or categories. | | [Pie](/data-analysis/development/runs/data-explorer/charts/pie-chart) | Circular graph that divides into slices, where the arc length of each slice is proportional to the quantity it represents. | | [Candlestick](/data-analysis/development/runs/data-explorer/charts/candlestick-chart) | Financial chart that displays how prices change for an asset over time, such as stocks and currency. | | [Map](/data-analysis/development/runs/data-explorer/charts/map-chart) | Uses a map to show how data is distributed across a geographic region. | | [Scatter](/data-analysis/development/runs/data-explorer/charts/scatter-chart) | Uses dots to show the relationship between two variables. | ## Filter conditions You can apply or remove filters to focus on specific parts of the visualized data. Filtering helps you explore subsets of the dataset without modifying or rerunning your pipeline. Filters apply only to the data currently loaded in the Data Explorer sample. FilterChart To apply a filter to your data: 1. Click **Filter** at the top of the chart. 2. Click **Add Filter**. 3. Configure your filter. Select the **Condition**, **Column**, and **Value** to filter for. 4. Click **Apply**. Click the trashcan icon at the top right corner of the filter to remove it. # Line chart Source: https://docs.prophecy.ai/data-analysis/development/runs/data-explorer/charts/line-chart Learn what parameters you need to configure a line chart A line chart visually represents data over time or along a continuous range. LineChart You can configure the following parameters for the chart: | Parameter | Description | | --------------- | -------------------------------------------------------------- | | X-axis column | Column used for the X-axis values, typically time or sequence. | | Y-axis column | Column with aggregated values used for the Y-axis. | | Min. Value | Minimum value displayed on the Y-axis. | | Max. Value | Maximum value displayed on the Y-axis. | | Tick Interval | Interval between tick marks on the Y-axis. | | Display Legends | Whether to display the legend on the chart. | | Enable Tooltips | Whether to display tooltips on hover. | | Show Grid Lines | Whether to display grid lines on the chart. | # Map chart Source: https://docs.prophecy.ai/data-analysis/development/runs/data-explorer/charts/map-chart Learn what parameters you need to configure a marker, displacement, and heat map chart A map chart plots geographic point data on a map. Prophecy supports three map chart types: a [marker map](#marker-map-chart), a [displacement map](#displacement-map-chart), and a [heat map](#heat-map-chart). A map chart option only appears when the dataset contains a column of geo points in WKT format, for example `POINT(-74.006 40.7128)`. Note that WKT orders coordinates as longitude, then latitude (this is the reverse of how coordinates are often read aloud). ## Marker map chart A marker map chart uses pins to represent individual data points on a map. For example, you might use a marker map to plot store locations, delivery addresses, or sensor sites. MarkerMapChart You can configure the following parameters for the chart: | Parameter | Description | Data type | Required | | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------- | -------- | | Column | Column containing geo points in WKT format. | WKT geo point | Yes | | Label Column | Column containing a name for each geo point, such as a `city_name` column. | Text | Yes | | Tooltip Columns | One or more columns whose values appear in the tooltip when a marker is clicked. | Any | No | | Fit Bounds | When enabled, zooms and pans the map to fit the extent of the plotted points, instead of showing the default world view. For example, if every point falls within the United States, enabling this shows only the United States. Off by default. | Boolean | No | Rows with a malformed WKT string, or with latitude/longitude values outside the valid range (latitude: -90 to 90, longitude: -180 to 180), are silently excluded from the map rather than raising an error. ## Displacement map chart A displacement map chart draws lines between origin and destination points to show movement or flows between locations. For example, you might use a displacement map to plot shipping routes, flight paths, or migration between cities. DisplacementMapChart Displacement map is only available as a map type when the dataset has more than one column of WKT geo points. You can configure the following parameters for the chart: | Parameter | Description | Data type | Required | | ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------- | ------------- | -------- | | Source Column | Column containing the starting geo point, in WKT format, for each path. | WKT geo point | Yes | | Destination Column | Column containing the ending geo point, in WKT format, for each path. | WKT geo point | Yes | | Source Label Column | Column containing an identifier for each starting point, such as a `city_name` column. | Text | Yes | | Destination Label Column | Column containing an identifier for each destination point, such as a `city_name` column. | Text | Yes | | Tooltip Columns | One or more columns whose values appear in the tooltip when a point or path is clicked. | Any | No | | Show Direction | When enabled, adds an arrowhead to each path indicating direction from source to destination. Off by default. | Boolean | No | | Fit Bounds | When enabled, zooms and pans the map to fit the extent of the plotted points instead of showing the default world view. Off by default. | Boolean | No | Rows with a malformed WKT string, or with latitude/longitude values outside the valid range, are silently excluded from the map rather than raising an error. ## Heat map chart A heat map chart shows density or intensity across a geographic area, weighted by a numeric value at each point. For example, you might use a heat map to plot where customer activity or order volume concentrates geographically. HeatMapChart You can configure the following parameters for the chart: | Parameter | Description | Data type | Required | | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------- | -------- | | Column | Column containing geo points in WKT format. | WKT geo point | Yes | | Value Column | Numeric column used to weight the intensity of the heat layer at each point. | Numeric | Yes | | Tooltip Columns | One or more columns whose values appear in the tooltip. | Any | No | | Radius | Controls the size, in pixels, of the blurred area drawn around each point. A larger radius produces bigger, more overlapping blobs and a smoother heat surface; a smaller radius produces tighter, more distinct hot spots. Defaults to 20. | Numeric | No | Radius is a fixed pixel size; it does not scale with zoom level. Zooming in or out changes the apparent size of each point's blob relative to the map, so a radius that looks right at a city-level zoom may cover a much larger area at a closer zoom, or blur many points together at a world view. Radius also determines the clickable/hoverable area used to trigger a point's tooltip. Rows with a malformed WKT string, or with latitude/longitude values outside the valid range, are silently excluded from the heat layer rather than raising an error. # Pie chart Source: https://docs.prophecy.ai/data-analysis/development/runs/data-explorer/charts/pie-chart Learn what parameters you need to configure a pie chart A pie chart is a circular graph that divides into slices, where the arc length of each slice is proportional to the quantity it represents. PieChart You can configure the following parameters for the chart: | Parameter | Description | | ------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | | Category column | Column used for pie chart segments. | | Y-axis column | Column containing values to aggregate into segment sizes. | | Chart angle | Starting angle in degrees for rendering the chart. | | Chart style | Style of chart, either pie or donut. | | Radius % (Pie style only) | Percentage of the chart canvas used for the pie's radius (the overall size of the pie chart). | | Inner and Outer Radius % (Donut style only) | Size of the inside and outside radius of the donut chart. | | Horizontal center % | Percentage from the left edge of the chart canvas for positioning the pie's center (the horizontal alignment of the chart). | | Vertical center % | Percentage from the top edge of the chart canvas for positioning the pie's center (the vertical alignment of the chart). | | Display Legends | Whether to display the legend on the chart. | | Enable Tooltips | Whether to display tooltips on hover. | # Scatter chart Source: https://docs.prophecy.ai/data-analysis/development/runs/data-explorer/charts/scatter-chart Learn what parameters you need to configure a scatter chart A scatter chart uses dots to show the relationship between two variables. ScatterChart You can configure the following parameters for the chart: | Parameter | Description | | --------------- | -------------------------------------------------------------------------------- | | X-axis config | Configuration for the X-axis, including min value, max value, and tick interval. | | Y-axis config | Configuration for the Y-axis, including min value, max value, and tick interval. | | Chart style | Style of the chart, either scatter or bubble. | | Display Legends | Whether to display the legend on the chart. | | Enable Tooltips | Whether to display tooltips on hover. | # Data exploration for Data Analysis Source: https://docs.prophecy.ai/data-analysis/development/runs/data-explorer/data-explorer Generate data samples through the pipeline during development During development, you can inspect data as it flows through your pipeline. Data Preview icons appear throughout the pipeline and let you open the Data Explorer. You can also display row counts directly on the canvas. Enable **Data Preview** from a gem's action menu to preload sample data after each run. Enable **Record Count** to display the number of rows produced by each gem. ## Understand icons 1. Data Preview icons appear on the pipeline canvas wherever data previews are available. By default, the icons are inactive row counts do not display (1). 2. If you enable **Data Preview**, an indicator appears on the gem and the Data Preview icon becomes active (2). 3. If you enable **Row Count**, a row count appears above the Data Preview icon (3). 4. Row counts also appear automatically at *model boundaries* (that is, immediately before a destination table). all features data sampling ### View Data Explorer To automatically load sample data during pipeline execution, choose Data Preview from the gem menu. The preview icon becomes active, and sample data is preloaded after each run. Click a Data Preview icon to open the Data Explorer. Data sample in a pipeline ### View complete dataset The Data Explorer loads a sample of your data by default. When you sort, filter, or search, these actions apply only to the visible rows in the sample. To work with the full dataset, do one of the following: * Click **Load More** at the bottom of the table until all rows are visible. * Click **Run** in the top-right corner of the preview. This refreshes the view and applies sorting and filtering to the entire dataset. ## Use the Data Explorer In the Data Explorer, you can: * Sort data by columns. * Filter rows by specific values. * Search across all values. * Show or hide columns. * Export the sample as a CSV or JSON file. * View your data as a chart in the [Visualization](/data-analysis/development/runs/data-explorer/charts/charts) tab DataExplorationSQL ## View record count To display the number of rows produced by a gem, choose **Record Count** from the gem menu. When you enable **Record Count**, Prophecy displays the row count above the Data Preview icon on the gem's output. This lets you quickly verify how many rows are flowing through your pipeline without opening the Data Explorer. Row counts are available as follows: * At model boundaries by default. * On intermediate gems when **Data Preview** is enabled. If you disable **Record Count**, any displayed row counts remain visible until you run the pipeline again or refresh the page. # Data profiling for Data Analysts Source: https://docs.prophecy.ai/data-analysis/development/runs/data-explorer/data-profile See statistics for data samples in your pipeline Data profiling allows you to view statistics on interim datasets in your pipeline. When you open a dataset's profile in the [Data Explorer](/data-analysis/development/runs/data-explorer/data-explorer), you can visualize value distributions and data completeness to ensure your data matches expectations. ## Quick profile The Data Explorer includes data profiles that are generated on your sample data. You'll be able to see high-level statistics for each column, including: * **Percent of non-blank values:** The percentage of values in the column that are not blank. * **Percent of null values:** The percentage of values in the column that are null. * **Percent of blank values:** The percentage of values in the column that are blank. * **Most common values:** Displays the top four most frequent values in the column, along with the percentage of occurrences for each. To view these statistics for your sample data, click **Profile** in the Data Explorer (1) to open the Data Profile view. Initially, Prophecy calculates Data Profile statistics using only the first 10,000 rows of data, but you can ask Prophecy to expand this calculation to the full data set. Before you do so, we recommend checking total row count by clicking the **Total Row Count** button at the bottom of the Data Explorer (2). Next, click **Load Full Profile** to calculate statistics for the entire dataset (3). Quick profile ## Expanded profile You can view a more detailed profile for any column in your dataset. When you open the expanded profile, Prophecy generates a deeper analysis of the column based on your sample data. All statistics in the expanded profile are computed on the same sample used in the Data Explorer. To view the expanded profile: 1. Click the dropdown arrow on the column you want to expand. 2. Select **Show Expanded Profile**. Show Expanded Profile When you load the expanded data profile, Prophecy generates a more in-depth analysis on the sample data. Expanded profile The expanded profile displays the following metrics: | Metric | Description | | ------------------------ | ----------------------------------------------------------------- | | **Data type** | Data type of the column. | | **Unique values** | Number of unique values in the column. | | **Longest value** | Longest value in the column and its length. | | **Shortest value** | Shortest value in the column and its length. | | **Most frequent value** | Most frequent value in the column and its number of occurrences. | | **Least frequent value** | Least frequent value in the column and its number of occurrences. | | **Minimum value** | Minimum value in the column. | | **Maximum value** | Maximum value in the column. | | **Average value length** | Average length of each value in the column. | | **Null values** | Percent and number of null values in the column. | | **Blank values** | Percent and number of blank values in the column. | | **Non-blank values** | Percent and number of non-blank values in the column. | | **Data summary** | Overview of the most common values in the column. | You can click between columns in the expanded profile for quick access. # Pipeline runs Source: https://docs.prophecy.ai/data-analysis/development/runs/execution Explore different ways you can run Prophecy pipelines The pipeline lifecycle consists of different stages, including development, production, and reporting. You can run your pipelines in different ways during each stage of the lifecycle. * **Development**: Press play in the canvas to run your pipeline and validate transformations by previewing data at each step. * **Production**: Use scheduled runs to automate pipeline execution at defined intervals. * **Reporting**: Run analysis dashboards to view and share pipeline results. This page explores these different run types in detail. ## Interactive runs in the canvas Prophecy lets you interactively run your pipeline in the pipeline canvas and preview data between each gem. This way, you can ensure that gems produce the expected output. There are two ways to start an interactive run: * **Click the large play button on the bottom of the pipeline canvas.** The whole pipeline runs. * **Click the play button on a gem.** All gems up to and including that gem run. This lets you test a small part of the pipeline, without consuming resources to run the whole pipeline. As gems run in your pipeline, sample outputs will appear after those gems. When you click on a data sample, Prophecy loads the data and opens the [Data Explorer](/data-analysis/development/runs/data-explorer/data-explorer). The Data Explorer lets you sort, filter, and search through the gem output. ## Scheduled runs Scheduling allows you to automate your data pipelines at predefined intervals. For each pipeline in your project, you can configure independent schedules that specify how often a pipeline runs and whether to send alerts during the automated runs. The execution environment of the scheduled run is determined during project publication. To learn more about deploying projects to specific execution environments, see [Versioning](/data-analysis/production/publication) and [Scheduling](/data-analysis/production/scheduling/scheduling). ## Executing pipelines via analysis dashboards You can also run pipelines with [analysis dashboards](/data-analysis/analysis/overview) in Prophecy. These dashboards enable non-technical users to run data pipelines through intuitive, form-based interfaces. By restricting access to pipelines themselves, you can provide proper guardrails for pipeline execution. ## External data handling Prophecy supports external sources and targets through [connections](/data-analysis/environment/connections/connections). Because SQL transformations require [tables](/data-analysis/gems/source-target/source-target), Prophecy Automate dynamically creates temporary tables in your SQL warehouse to process this data. Temporary tables act as intermediaries that allow external data to be processed using SQL logic. In other words, they enable dbt and SQL to transform external data as if it were native to the warehouse. These tables use ephemeral materialization, which means they exist only during query execution and do not appear in your warehouse or pipeline canvas. ## Where pipelines run All pipeline runsβ€”whether interactive, scheduled, or triggered through appsβ€”execute using: * **SQL warehouse**: Processes SQL transformations using [dbt](https://docs.getdbt.com/docs/build/models) * **Prophecy Automate**: Extends SQL warehouse capabilities by providing orchestration and ingress/egress features Different components of the same pipeline can run in different places. SQL transformations execute in your SQL warehouse, while orchestration, data ingestion, and data egress operations run through Prophecy Automate. Your [Prophecy fabric](/data-analysis/environment/fabrics/prophecy-fabrics/) configuration determines the compute resources, connection details, and runtime settings that apply to all run types. # Runtime logs Source: https://docs.prophecy.ai/data-analysis/development/runs/runtime-logs View live and historical information about each operation performed during a pipeline run Runtime logs provide visibility into pipeline execution in Prophecy. They help you understand how a pipeline ran, which steps succeeded or failed, how long operations took, and where errors occurred. Prophecy supports both: * **Live runtime logs** for active or recent runs in the project editor. * **Historical runtime logs** for completed runs, which are snapshots of the pipeline at execution time accessed through the Observability interface. ## Open runtime logs You can open runtime logs from the information footer at the bottom of the project editor. Runtime logs update in real time while your pipeline runs. Each log entry provides structured details that help you trace execution and diagnose issues. Logs include: * **Pipeline steps**: Execution details for every processing step, mapped directly to the gems in your visual pipeline. * **Execution sequence**: The exact order in which steps were carried out. * **Timestamp**: The start time of each step. * **Duration**: How long each step took. * **Status information**: Success, warning, and failure details for each operation. When an error occurs, the message usually comes directly from the execution engine. For example, if your pipeline is running on a Databricks SQL warehouse, errors originate from Databricks. If you're using [Prophecy In Memory](/data-analysis/environment/fabrics/prophecy-fabrics), errors come from DuckDB (the underlying warehouse). For deeper troubleshooting, consult the corresponding warehouse documentation to interpret execution-specific error details. You can also see runtime log information by hovering over each gem in the pipeline canvas. This shows the duration and success/failure status of that step in the most recent run. Use this to quickly spot bottlenecks, failed operations, or unusually slow steps without opening the full log view. ## Use runtime logs In the Runtime Logs tab, you can use several tools to help navigate and analyze execution details. 1. **Filter logs by keywords**: Find specific information by entering search terms 2. **Sort logs**: Organize logs chronologically 3. **Filter logs by log level**: Focus on info, warning, or error messages 4. **Download logs**: Save logs to your machine as a text file 5. **Expand log screen**: Display a larger view of your logs 6. **Clear all logs**: Remove log information for previous runs Closing Runtime Logs does not immediately clear log information for recent runs. Recent runtime logs may be restored automatically while they remain available in persistent storage. Runtime logs If your pipeline fails to run, you'll see an error message explaining the reason for the failure displayed at the top of the corresponding log. ## Access historical runtime logs You can reopen runtime logs for completed pipeline runs through [the Observability interface](/data-analysis/production/monitoring). Historical runtime logs help you investigate failures, compare past executions, validate pipeline behavior, and understand how a pipeline operated at a specific point in time. To access historical runtime logs: 1. Click **Observability** in the left sidebar and open the **Run History** tab. 2. Locate the pipeline run you want to inspect and click **See logs** in the far-right column. Access historical runtime logs Prophecy opens the selected pipeline run in **Historical Mode**, showing: * Runtime logs for that execution. * The pipeline snapshot associated with the run. * Historical component execution states and statuses. ## Historical Mode When you open historical runtime logs, the pipeline opens in **Historical Mode**. Historical Mode reconstructs the pipeline exactly as it existed during the selected execution. This includes the pipeline graph, runtime logs, and component execution statuses. Historical runtime logs view In Historical Mode: * The pipeline canvas becomes read-only. * Runtime logs are restored for the selected run. * Components display their historical execution states. * Execution progress and failure states are preserved. * You can switch between historical runs to compare executions. * You can rerun historical pipeline snapshots. A banner at the top of the canvas indicates that you are viewing a historical pipeline snapshot. From this banner, you can: * Return to overall run history. * Exit Historical Mode and return to the latest pipeline version. ## Pipeline snapshots and retention Historical runtime logs rely on persisted pipeline snapshots and execution metadata. When available, Prophecy stores: * Runtime log output. * Pipeline execution state. * Component execution progress and statuses. * Pipeline snapshots associated with completed runs. These artifacts are retained for a limited period of time and may eventually expire depending on system configuration and storage limits. Not every historical run may have an available pipeline snapshot. If a snapshot cannot be restored, Prophecy displays a message indicating that the historical pipeline snapshot could not be found and provides an option to return to the latest pipeline version. When loading a historical run, Prophecy displays a loading state while runtime logs and pipeline snapshots are restored. ## SQL warehouse logs You can find additional logs specific to pipeline queries in your SQL warehouse. If you are connected to a Databricks SQL warehouse, you can find these logs directly in Databricks. This extra visibility into query execution in Databricks helps you troubleshoot pipeline behavior and improve performance. If you have access to your Databricks workspace: 1. Navigate to **SQL > Query History** in the Databricks sidebar. 2. Locate the SQL queries run during your pipeline's execution window. 3. Click on a query to see: * Full SQL query * Status (Success or Error) * Start and end times * Error messages * Execution plan or profile ## What's next Now that you know about runtime and publish logs, you might want to learn about [audit logs](/data-analysis/administration/management/audit-logs) and [system logs](/data-analysis/administration/getting-help/prophecy-details), which provide visibility into platform-level and administrative activity. * **Audit logs** track events across entities like projects, pipelines, jobs, and fabrics. These logs help administrators understand who did what and when. For example, these logs can capture who creates a pipeline, releases a project, modifies a dataset, etc. * **System logs** capture backend infrastructure details to support platform monitoring and troubleshooting. These include Kubernetes cluster configuration (such as resource quotas and node settings), cluster custom resources, config maps and files, and resource consumption metrics. Unlike runtime logs, which focus on the execution of a specific pipeline run, audit and system logs provide a broader view of how users and the platform are operating over time. # Canvas annotations Source: https://docs.prophecy.ai/data-analysis/development/studio/canvas-annotations Leave comments on your project canvas to annotate pipelines Prophecy enhances pipeline readability and comprehension through several annotation features: * **Gem comments:** Copilot generates comments and documentation for individual pipeline components (gems), providing context and explaining complex transformations. You can also create and update gem comments yourself. * **Canvas annotations:** Add free-form text annotations directly on the pipeline canvas, highlighting key steps, providing explanations, or documenting assumptions. * **Gem labels and icons:** Labels and icons on each gem allow users to visually categorize and identify pipeline components, improving overall pipeline clarity and organization. In many cases, pipelines can become quite large to accommodate complex transformation requirements. Learn how to add canvas annotations in the following sections. ## Add an annotation To add an annotation to a pipeline: 1. Open a pipeline in a project. 2. Click on the annotate button in the bottom left corner of the canvas. 3. Drag the text box to your desired location in the canvas. 4. Add your own text or image to the annotation. 5. Format the text using the formatting toolbar. ## Example Let's say you have a pipeline that ingests customer data, cleans it, and then applies transformations before loading it into a database. To help your team understand the different stages, you can add annotations like this: * Add an annotation near the input source stating: `Ingesting raw CSV files from S3 bucket` * Annotate above the transformation gems with: `Removing duplicates and normalizing column names` * Mark next to a Macro gem: `Using imported dbt Pivot macro` * Add an annotation near the last node saying: `Writing transformed data to Databricks catalog` This helps your team quickly grasp what each part of the pipeline does without digging into the details of every transformation. # Use visual containers Source: https://docs.prophecy.ai/data-analysis/development/studio/containers Organize your pipeline into containers to group related transformations Containers allow you to divide your pipeline into logical sections. They help you visually organize related gems, making pipelines easier to read and understand. You can also collapse and expand containers to make room on the canvas and differentiate containers by color. For example, you can use containers to group transformations by: * Pipeline stage: Separate input preparation, transformation logic, and output writing into distinct sections. * Gem phase: Group transformations that execute together to reflect the sequential flow of the pipeline. * Input source: Organize steps that prepare data from different sources before they are joined downstream. Pipeline grouped into two visual containers ## Create a container To create a container: 1. Drag to select multiple gems on the canvas. 2. At the bottom of the canvas, click **Actions**. 3. Select **Group**. The selected gems are added to a new container, named `VisualGroup_1` by default. ## Manage containers Containers have their own **...** menu, separate from the action menu on individual gems. To manage a container, click its **β‹―** menu: * **Ungroup** β€” Remove the group and keep its gems on the canvas. * **Add to AI** β€” Add a mention for the container to Chat. * **Explain** β€” Copilot provides an explanation of what the container groups. * **Attach Annotation** β€” Add an expanded label for the container. * **Edit Background Color** β€” Change the container's background color to visually distinguish it from others. * **Collapse** β€” Hide the container's gems and simplify the canvas. (You can reverse this using an **Expand** command.) Pipeline with collapsed container You can also: * **Rename a container** β€” Click the container name and enter a new one to clarify its purpose. * **Move gems in or out of a container** β€” Drag the gem in and out across the container boundary. * **Edit container size** β€” Drag the container edges or corners to resize. # Development settings Source: https://docs.prophecy.ai/data-analysis/development/studio/development-settings Configure development settings for your project Development settings control how Prophecy processes your project during development. To access these settings, open the project settings menu beside your project's name in the Studio header and select **Development Settings**. You'll see the following settings: * [Variant interference data sampling limit](#variant-interference-data-sampling-limit) * [Backend language](#backend-language) * [Use original project browser layout](#use-original-project-browser-layout) * [Show models](#show-models) * [Enable manual compilation](#enable-manual-compilation) ## Variant interference data sampling limit This number controls the maximum number of records Prophecy parses to infer the schema of variant data types. The default limit is 100 records. Adjust this value based on your data characteristics: * **Increase the limit** for complex structures where 100 records may not capture all schema variations. * **Decrease the limit** for simple datasets where processing fewer records improves performance while maintaining schema accuracy. This setting applies globally to all variant schema inference operations in your project. ## Backend language Private Preview This setting specifies the programming language used for project compilation and code generation. Currently, Prophecy supports **SQL** and **Python** as the backend language for data analysis projects. The backend language determines: * The syntax used in generated code * The runtime environment requirements * The project artifacts produced You can switch between the backend languages using the dropdown menu. When you switch your project's backend language (like from SQL to Python or vice versa), not every feature or configuration translates perfectly. For example: * Unit tests are only available for PySpark but not for SQL. * Some gems may not have an equivalent in the other language. If a gem can't be converted, you'll see an error stating that Prophecy does not support the gem in the current language. * Write modes can vary between languages. For example, `SCD2 Soft Delete` is only available for Python. * Macros in SQL do not have a direct Python equivalent, so switching may make some logic untranslatable or unavailable. Prophecy is actively developing these features and is aiming for full parity between languages in upcoming releases. As a best practice, **always save or commit your work before switching** so you can easily revert if something doesn't translate. ## Use original project browser layout When you open a project, Prophecy restores the last saved project state when one is available. If there is no saved state, the default layout depends on whether AI is enabled and whether the project already contains entities. You can reverse this layout by clicking the **Use Original Project Browser Layout** toggle. ## Show models Many pipeline transformations are compiled into models under the hood. If you are working on a pipeline, you can view and edit the code of underlying dbt models in a pipeline. (You cannot visually edit these underlying models). To view the dbt models created in the backend for pipelines in SQL projects, click the **Show Models** toggle. ## Enable manual compilation To enable this feature, you must set `EDITOR_ON_DEMAND_COMPILATION: 'true'` in `cp-env-cm` of your Prophecy deployment. This toggle controls whether Prophecy compiles your project automatically as you update your project. * **Disabled** (default): Prophecy compiles automatically. This provides immediate feedback on compilation errors and ensures your project remains in a functioning state. * **Enabled**: You compile the pipeline or project manually by clicking **Compile** in the Studio footer. Use manual compilation when working with large projects where automatic compilation slows down your workflow or any time you need more control over when compilation occurs. Manual compilation With manual compilation enabled, your project may contain compilation errors that won't be detected until you compile manually. Ensure you compile before running pipelines or deploying your project. # Pipeline folders Source: https://docs.prophecy.ai/data-analysis/development/studio/folders Keep your project organized by grouping pipelines in folders Use folders to organize pipelines within your Prophecy project and facilitate project navigation and collaboration. Folders help you manage large numbers of pipelines by grouping them systematically. You'll be able to see new folders in both the visual and code view of the project. Pipeline folders must be created inside the `pipelines` directory of the project. ## Limitations Pipeline folders have the following limitations: * Folders can only be created, not edited or removed * You cannot rename folders after creation * You cannot move pipelines between folders after they are created ## Create a folder ### Method 1: Create the folder directly This is the most straightforward method to create a new folder. 1. Open the project editor. 2. In the left sidebar, hover over **Pipelines**. 3. Click the **+** (plus) icon. 4. Click the folder icon next to the directory path. 5. Select the `pipelines` directory. 6. Click **Add Folder**. 7. Enter your folder name. 8. Click **Save**. The folder is created. You can either create a new pipeline inside the folder or exit the new pipeline dialog. If you exit the dialog, you can add a pipeline to the new folder later. ### Method 2: Create the folder and a new pipeline For this method, you type the pipeline path directly. However, it requires that you also create a new pipeline simultaneously. 1. Open the project editor. 2. In the left sidebar, hover over **Pipelines**. 3. Click the **+** (plus) icon. 4. In the **Pipeline Name** field, name your new pipeline. 5. In the **Directory Path** field, type the new folder path (for example, `pipelines/finance`). 6. Click **Create** to save the new pipeline and the new folder. # Status messages Source: https://docs.prophecy.ai/data-analysis/development/studio/messages Understand messages that appear in the Studio Status messages appear in Studio to indicate what Prophecy is doing behind the scenes as you make changes or run your pipeline. You'll see these messages in the bottom-right corner or the center of the Studio canvas. ## Reference You'll see the following messages in Studio. | Message | Description | | ------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Initializing project** | Prophecy is loading project files and setting up the Studio environment. This message blocks interaction with the canvas until initialization completes. | | **Processing changes** | Prophecy has received your edits and is checking for side effects that need to be handled. The system identifies dependencies and updates related project files automatically.

For example, adding a table from the Environment tab to the pipeline canvas also creates an entry in `sources.yml`. | | **Generating code** | Prophecy is converting your visual pipeline into SQL files. During this process, Prophecy:
  • Scans all gems in the current pipeline for changes.
  • Generates SQL code for each gem.
  • For custom macro gems, executes the `apply` method you implemented in a separate sandbox environment.
Code generation runs automatically when you modify gems. Duration depends on the number of gems and complexity of transformations in your pipeline. | | **Compiling** | Prophecy is performing checks and computing output schemas for each gem's output. Compilation happens automatically after code generation. | | **DBT initialization** | Prophecy is preparing the [dbt environment](https://docs.getdbt.com/docs/introduction) for your project (since Prophecy uses dbt behind the scenes for transformations). Before running pipelines, Prophecy:
  • Creates a sandbox environment loaded with all project files.
  • Updates the sandbox with recent changes.
  • Compiles models using dbt to produce final runnable SQL code.
This message appears when you start a pipeline run. | | **Running** | Prophecy is executing your pipeline. This includes orchestrating the run, managing dependencies, and coordinating execution across different components. | | **Running query in warehouse** | Prophecy is executing SQL queries directly in your data warehouse. Different parts of your pipeline run in different environments:
  • SQL transformations execute in your SQL warehouse.
  • Orchestration, data ingestion, and data egress operations run through Prophecy Automate.
| ## Duration Status messages may remain visible longer in large projects because: * More gems require more code generation. * Complex dependencies need more processing time. * Larger schemas take longer to compute. * More files need to be synchronized with the dbt sandbox. If a message persists for an unusually long time, try refreshing the page. If the problem persists, [contact support](/data-analysis/administration/getting-help/get-in-touch). # SQL Shell Source: https://docs.prophecy.ai/data-analysis/development/studio/sql-shell Run ad-hoc SQL queries directly against the warehouse configured for your project. The SQL Shell provides an interactive query interface inside the SQL IDE. You can use it to run ad-hoc SQL queries against the warehouse configured for the current project without leaving the pipeline or analysis context. The SQL Shell opens as a separate tab alongside your existing tabs and includes a SQL editor, query results panel, query history, and query execution controls. ## Run query 1. Click the **...** menu in the upper-right corner of the pipeline canvas or SQL IDE header. 2. Select **SQL Shell**. 3. Wait for the shell to initialize and open in a new tab. 4. Enter a SQL query in the editor. 5. Click **Run**. 6. Review query results in the lower panel. You can cancel a running query using the **Cancel** button. The SQL Shell sends the query text directly to the backend with only whitespace trimming applied by the frontend. Query execution behavior, including support for multi-statement queries or DDL statements, depends on the connected fabric and backend configuration. ## View query results Query results appear in the lower results table with paginated loading. Results are returned in pages of 100 rows at a time. To retrieve additional rows, load more results from the table interface. The shell streams execution output live while queries run. Timing information, logs, and row counts appear as execution progresses. ## View query history Click the **History** tab to view previously executed queries. Query history is initialized when the shell session starts and persists for the duration of the connection. If the page refreshes or the connection reconnects, the history reloads from the backend. # Connection behavior The SQL Shell uses the warehouse and credentials configured for the project's fabric. You do not select or enter credentials directly in the shell. The shell maintains a persistent WebSocket connection during the session. This enables live query output and automatic reconnection behavior if the connection temporarily drops. Enterprise network proxies or firewalls that block secure WebSocket (`wss://`) connections may prevent the SQL Shell from functioning correctly. # Availability and limitations The SQL Shell is available only in the SQL IDE for SQL Pipelines and Analyses/Apps. The first time you open the SQL Shell for a fabric, there may be a startup delay while the backend creates a command executor, opens a JDBC connection to the warehouse, and validates the connection. If initialization fails, the shell displays an error before opening. Common causes include invalid credentials, unavailable warehouses, paused warehouses, or incorrect connection configuration. The SQL Shell is disabled in the following situations: * When viewing a pipeline in historical mode. * When the environment is in read-only mode. # Troubleshooting Common failure scenarios include: * Fabric mis-configuration, such as invalid credentials, missing warehouses, or insufficient permissions. * Network or connectivity issues that block the WebSocket connection. * Paused or unavailable warehouses. Resolve fabric configuration issues at the fabric level. For connectivity issues, contact your administrator. # Studio interface Source: https://docs.prophecy.ai/data-analysis/development/studio/studio Learn how to use Prophecy's built-in IDE The Studio is Prophecy's built-in IDE for building [pipelines](/data-analysis/development/pipelines/data-analysis-pipelines). This page walks through each area of the interface, from the sidebars and canvas to the header and footer. ## Sidebars When you open Studio, Prophecy opens the last saved state of the project when one is available. If there is no saved state, Prophecy opens a project landing page where you can open existing entities or create new ones. New projects without entities open with Agent chat maximized when v4 AI is enabled. ### Agent chat Use the chat interface to interact with Prophecy Agents. Learn more about how to use AI for your project in the [Agent Overview](/data-analysis/ai/agent/agent) documentation. When Agent chat is maximized, it appears full screen. You can minimize the chat to show the canvas landing page. You can also rename a chat by clicking the chat name, and manage chats from chat history. See [using chat](/data-analysis/ai/agent/chat/using-chat) for more details. ### Project browser In the Visual view, browse through project entities such as pipelines, tables, analyses, functions, and more. When you switch to the Code view, you will see the project repository directory instead. Opening or closing the Project Browser only changes the Project Browser panel. It does not control whether the landing page or canvas is visible. #### Environment browser The **Environment** tab lets you access data directly in Prophecy from [connections](/data-analysis/environment/connections/connections) defined in your attached fabric. You can: * Browse or search for available datasets * Drag datasets directly onto the visual canvas * Add new connections to the attached fabric On the **Free** and **Professional** Editions, the Prophecy Warehouse connection has additional functionality. Next to each table, click the `...` ellipses menu to access the following options: * **Preview**: See a preview of the data in the selected table. * **Share**: Invite users to your team so they can access the fabric's data. * **Rename**: Update the name of the stored table. * **Delete**: Delete the stored table. When you update the name of a table, any gems pointing to this table location are not automatically updated; you must update them yourself with the new name. ## Pipeline canvas The pipeline canvas is the workspace where you can add and connect various gems to build your pipeline. It provides a drag-and-drop interface for designing your data flow. ### Gem drawer At the top of the pipeline canvas, the **gem drawer** displays gem categories such as Transform and Join, which contain all the gems available for use in your pipeline. You can bookmark gems by clicking the star to their right in the gem dropdown menus. Bookmarked gems then appear in a Favorites bar in the Gem drawer. To learn more, visit [Gems for Data Analysis](/data-analysis/gems/gems). ### Annotate button At the bottom left of the canvas, the annotate button lets you add free-form text annotations directly on the pipeline canvas. To learn more, visit [Canvas annotations](/data-analysis/development/studio/canvas-annotations). ### Group gems into containers Drag to select multiple gems, then use the **Actions** menu at the bottom of the canvas to group them into a container. Containers help you visually organize related gems on the canvas. To learn more, visit [Visual containers](/data-analysis/development/studio/containers). ### Alt-key row count and connection reveal Holding **Alt** (**Option** on Mac) on the canvas switches interim row-count labels from a compact, rounded format to the exact count. For example, `100K rows` becomes `100,000 rows`, and `10.2K rows` becomes `10,166 rows`. Releasing **Alt** reverts to the compact display. This applies to: * SQL IDE canvas interim icons * Spark/Workflow canvas edge labels * The interim **view rows** action button Holding **Alt** also reveals all [wireless connections](/data-analysis/development/studio/wireless-connections) on the canvas at once. Normally, a wireless connection is only visible while its edge is selected. Releasing **Alt** hides them again; any already-selected edges stay visible regardless. Counts under 1,000 (for example, `5 rows`) look the same in both compact and exact format. To see the row-count behavior, run the pipeline first so interims have row counts to format. ### Undo and redo Use `Cmd+Z` (macOS) or `Ctrl+Z` (Windows/Linux) to undo your most recent change on the canvas, and `Cmd+Y` / `Ctrl+Y` to redo a change. This covers actions such as: * Gem property changes. * Form field edits. * Opening and closing dialogs. * Switching tabs. * URL search parameter changes. * Navigating between subgraphs or components. You can undo up to approximately 50 recent actions. Undo and redo primarily support changes to your pipeline graph. Some actions fall outside this scope and can't be undone, including: * Creating, opening, or deleting a pipeline. * Creating or deleting a dataset. * Project-level actions like scheduling, committing, or publishing. When you perform one of these irreversible actions, Prophecy warns you that the action cannot be done. Undo and redo are disabled during onboarding tours. ### Run button The **run** button triggers pipeline execution. This allows you to test and run the pipeline in real-time, which makes it easier to troubleshoot and verify the pipeline's performance before deployment. To learn more, visit [Pipeline execution](/data-analysis/development/runs/execution). ### Pipeline header In the subheader on the pipeline canvas, you can switch between pipeline and analysis tabs, as well as access a pipeline-scoped `...` dropdown menu. Here, you can schedule the pipeline; open pipeline parameters; [rename, duplicate, or delete the pipeline](/data-analysis/development/pipelines/data-analysis-pipelines); and access the [SQL shell](/data-analysis/development/studio/sql-shell). ## Header The Studio header includes the following elements. ### Project dropdown menu You can access various project settings, dependencies, and metadata from the project settings `...` dropdown menu beside the project name. You can also clone or delete the project from this menu, as well as export compiled code, view the project's version history, and view other project details. Clicking the project icon or project name opens the project landing page in the canvas. ### Parameters Each project has a small folder icon in the header. It will say **default** by default. This shows you the active project parameter set for the project. To access both [project and pipeline parameters](/data-analysis/development/parameters/parameters), click the folder icon. ### Doc-Visual-Code toggle Switch from the Visual view to the Code view to see your visual pipeline compiled into code. This view helps users who prefer working with code to understand the underlying logic of gems and pipelines. If you have [generated documentation for your pipeline](/data-analysis/ai/agent/documentation/documentation), you can view it by clicking the **Docs** tab. ### Fabric status This icon shows you the current status of your fabric connection. When you're connected to a fabric, it will show a green dot. When you're not connected, it will be gray. Click the icon to connect or disconnect from a fabric. ### Version menu If you create your project using the [simple Git storage model](/data-analysis/development/versioning/version-control), you will see the version menu in the project header. Use this menu to save your project, publish your project, or view your project history. If you create your project using the [normal Git storage model](/data-analysis/development/versioning/version-control), you will see the Git workflow in the project footer. Open the Git workflow to perform actions like committing, merging, or deploying the project. ## Footer The Studio footer includes the following elements. ### Problems panel The Problems panel highlights any issues or errors in your pipeline that need attention. It provides detailed feedback on what needs to be fixed to ensure that your pipeline runs successfully. ### Runtime Logs [Runtime logs](/data-analysis/development/runs/runtime-logs) offer detailed insights into the status and progress of your pipeline executions. They provide a step-by-step trace of how each transformation or action was performed, any errors, and other progress messages. ### Git workflow If you create your project using the [normal Git storage model](/data-analysis/development/versioning/version-control), you will see the Git workflow in the project footer. Open the Git workflow to perform actions like committing, merging, or deploying the project. # Wireless connections Source: https://docs.prophecy.ai/data-analysis/development/studio/wireless-connections Simplify complex pipelines by hiding visual connections while preserving execution logic Wireless connections let you hide visual lines between gems on the canvas while retaining the logical connection between gems. This option lets you reduce visual clutter, making large or evolving pipelines easier to read. Wireless connections affect only the canvas presentation. They do not change execution order, data dependencies, or runtime behavior. ## Make connections wireless You can convert wired connections into wireless ones from a gem's [action menu](https://docs.prophecy.ai/data-analysis/gems/gems#action-menu). To make a connection wireless: 1. Click the gem whose connection you want to modify. 2. Open the gem's action menu (`...`). 3. Select one of the following options: * **Make incoming connection wireless**. * **Make outgoing connection wireless**. The visual line disappears, but the data flow between the gems is preserved. Wireless connection applied For gems with multiple incoming or outgoing connections, you can make all connections wireless by choosing a **Make incoming connections wireless** or **Make outgoing connections wireless** option. ## Restore wired connections To restore visual connections: 1. Click the gem with wireless connections. 2. Open the gem's action menu (`...`). 3. Select one of the following options: * **Make incoming connection wired**. * **Make outgoing connection wired**. The visual connection lines are restored. Before wireless connection applied ## When to use wireless connections Wireless connections are most effective when: * Pipelines contain many parallel branches or long cross-canvas connections. * You want to reorganize or refactor a pipeline without rearranging large sections of the canvas. * Visual clarity is more important than showing every physical connection. # Project tests Source: https://docs.prophecy.ai/data-analysis/development/tests/project-tests Create a pipeline that validates specific data conditions Project tests are custom SQL queries that validate data conditions. You build each test as a visual pipeline ending with a Data Test gem. The pipeline returns data that the Data Test gem evaluates using configurable parameters to determine whether the test passes or fails. Use project tests to validate data that spans multiple models or tables, or to test specific transformation logic. Common use cases include verifying referential integrity across related tables, ensuring aggregated values meet business rules, or validating that complex transformations produce expected results. Project tests are based on [dbt singular data tests](https://docs.getdbt.com/docs/build/data-tests#singular-data-tests). Use **project tests** when you need to test a specific workflow or combination of models that won't be reused elsewhere. Use [test definitions](/data-analysis/development/tests/table-tests) when you want to apply the same standardized test (like checking for uniqueness or null values) across multiple tables or models in your project. ## Understand project test flow Project tests evaluate data through a two-step process: 1. **Pipeline execution**: Your visual pipeline executes and returns data. The pipeline can filter, transform, or aggregate data to prepare it for evaluation. 2. **Data Test evaluation**: The Data Test gem applies the **Failure Calculation** to the pipeline output, then evaluates the **Error If** and **Warning If** conditions against that calculated value to determine the test result. The test passes when the **Error If** and **Warning If** conditions are not met. ## Data Test gem parameters When you create a project test, a Data Test gem appears on an otherwise empty canvas. The Data Test gem evaluates the rows returned by your pipeline and determines whether the test passes or fails based on the following parameters: | Parameter | Description | Default | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------- | | **Failure Calculation** | Expression that calculates a value from the test results. The **Error If** and **Warning If** conditions evaluate this calculated value. | `count(*)` | | **Limit** | Maximum number of failure rows to return. Set this to reduce query execution time and resource usage when testing large datasets. The test stops after finding the specified number of violations. | Empty (no limit) | | **Severity** | Determines whether failures return an error or warning. When set to `error`, the test checks **Error If** conditions first, then **Warning If** conditions. When set to `warning`, only **Warning If** conditions are checked. | `error` | | **Error If** | Condition that triggers an error. Evaluates the **Failure Calculation** result. If true and severity is `error`, the test fails with an error status. | `!=0` | | **Warning If** | Condition that triggers a warning. Evaluates the **Failure Calculation** result. If true, the test returns a warning status (regardless of severity setting). | `!=0` | If you leave all parameters empty, Prophecy uses the default values. Customize parameters when you need more control over test behavior. The Data Test gem evaluates your pipeline output in this order: 1. Applies the **Failure Calculation** expression, producing a single value. 2. Evaluates the **Error If** condition against the calculated value (if severity is `error`). 3. If the **Error If** condition is not met, evaluates the **Warning If** condition. 4. Returns the test result. | Example | Failure Calculation | Error If | Description | | ------------------- | ------------------- | -------- | ------------------------------------------------------------------------ | | **Default** | `count(*)` | `!=0` | Counts rows. Fails if any rows exist, passes if no rows. | | **Threshold-based** | `count(*)` | `>100` | Counts rows. Fails only if more than 100 rows exist. | | **Aggregate-based** | `sum(amount)` | `<0` | Sums a column. Fails if the sum is negative, passes if zero or positive. | ## Build and run a test The following example procedure creates a test named `assert_total_payment_amount_is_positive` that validates payment data. For each order, there might be multiple transactions, where negative transactions represent refunds. The total payment amount for an order is the sum of all transaction amounts and should never be negative. If we use the default parameters, we need to build a pipeline that: 1. Aggregates payments by order to calculate total amounts 2. Filters to return only orders with negative totals (violations) 3. Passes these violation rows to the Data Test gem With default parameters `count(*)` and `Error If !=0`, if any orders have negative totals, the Data Test gem receives rows, the count is greater than 0, and the test fails. If no orders have negative totals, the count is 0, and the test passes. ### 1. Create a project test entity To develop a project test, start by opening a project: 1. In the left sidebar, click **+ Add Entity**. 2. Hover the **Tests** option and select **Project tests**. 3. Enter a name for your test, such as `assert_total_payment_amount_is_positive`. 4. Keep the default path `tests` where Prophecy will store the test. 5. Click **Create**. ### 2. Build the test pipeline Build a pipeline that returns data for the Data Test gem to evaluate. In this example, we'll return only rows that violate the business rule, which works well with default parameters. 1. Drag a table onto the canvas. In this example, we'll use a `payments` table: | order\_id | amount | | --------- | ------- | | ORD-001 | 150.00 | | ORD-001 | -25.00 | | ORD-002 | 200.00 | | ORD-003 | 100.00 | | ORD-003 | -150.00 | 2. Add an **Aggregate** gem after the `payments` table. 3. Configure the Aggregate gem to sum the `amount` column grouped by `order_id`. 4. Add a **Filter** gem after the Aggregate gem. 5. Configure the Filter gem with the condition `amount_sum < 0` to return only orders where the total amount is negative. 6. Connect the Filter gem to the **Data Test** gem. Because `ORD-003` has a negative total amount `(100.00 + (-150.00) = -50.00)`, the test will return an error. ### 3. Review the Data Test gem Since we are using the default parameters, we don't need to configure the Data Test gem. However, we can review the default parameters and the SQL query to understand how the test works. 1. Open the **Data Test** gem. 2. Keep the default parameters. Without modifying the parameters, the test returns an error if the input to the Data Test gem has any rows. 3. Review the **Final Query** code editor. This displays the SQL query generated from your visual pipeline. The test executes this query against your data warehouse. You don't need to edit it, but reviewing it helps verify the test logic. 4. Click **Save**. ### 4. Run the project test Run the whole pipeline to see the test result. 1. Click the **Play** button on the canvas or on the Data Test gem. 2. The test executes the SQL query from the **Final Query** against your data warehouse. 3. Click **See Run Details** in the top right of the canvas to view the test summary. See Run Details 4. Review the test status: succeeded, warning, or failed. 5. For failed tests, expand the logs section to view detailed dbt execution logs. The logs show the SQL query that was executed and the test result. Test logs You can execute a partial pipeline run by clicking play on an intermediate gem. However, since the execution stops before reaching the Data Test gem, the test will not run. ### 5. Schedule test runs Project tests can run as part of a pipeline schedule. 1. Open the pipeline you want to associate with project tests. 2. In the project header, click **... > Schedule**. 3. Edit the existing schedule or configure a new schedule. 4. Under **Project level tests**, select the tests you want to run. 5. Click **Confirm** to save the changes. If a scheduled test fails, you'll be able to see the test logs in the [Observability](/data-analysis/production/monitoring) interface. You must enable the schedule and publish the project to activate the automation. Learn more in [Schedule activation](/data-analysis/production/scheduling/scheduling#schedule-activation). When scheduling models only (not pipelines), configure tests through jobs: 1. In the left sidebar, click **+ Add Entity** > **Job**. 2. Enter a name for your job and click **Create New**. 3. Drag a **Model** gem to the canvas. 4. Click the model to open its properties. 5. Select the database object to test. 6. Select the **Run tests** checkbox in the left sidebar. 7. Verify your **project**, **model**, and **fabric** settings. 8. Click **Save**. ## Troubleshooting test failures When a project test fails, determine whether the failure indicates a data quality issue or a configuration problem. ### Test failure (data quality issue) The test evaluates the pipeline output using the configured **Failure Calculation** and **Error If** or **Warning If** conditions, and the conditions are met, indicating that your data violates the test criteria. The test is working correctly, but your data doesn't meet the expected criteria. When this happens, you should review the test logs to see which conditions were met. Either fix the data issue or consider adjusting the **Error If** or **Warning If** thresholds if the current values are too strict or too lenient. For example, in the `assert_total_payment_amount_is_positive` test with default parameters, if the query returns rows and the count is greater than 0, it means there are orders with negative total payment amounts. You would need to investigate why those orders have negative totals. ### Execution error (configuration or setup issue) The test cannot run properly due to a technical problem. Common causes include: * The input table no longer exists or the input data sources are inaccessible. * The **Failure Calculation** function is invalid or contains syntax errors. * The **Error If** or **Warning If** conditions are invalid or contain syntax errors. * The SQL query itself has syntax errors. # Table tests Source: https://docs.prophecy.ai/data-analysis/development/tests/table-tests Create reusable data tests using parameterized SQL Table tests are reusable SQL queries that validate your data quality. You write a test definition once, then apply it to any table or model in your project to check for data problems. Use test definitions to catch data issues before they affect your analysis or reports. For example, you might create a test to verify that customer IDs are unique, that required fields aren't missing, or that two related tables have matching values. When you run these tests, you'll be able to identify any problems in your data. Table tests are based on [dbt generic data tests](https://docs.getdbt.com/docs/build/data-tests#generic-data-tests) behind the scenes. When you add a test definition to a table, Prophecy adds the test as a property of the corresponding dbt model. SQL fabrics configured with BigQuery and a CMEK are not compatible with data tests. ## Default test definitions Prophecy includes a set of default test definitions that you can use to get started. | Test definition | Description | | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | | **Unique** | Validates that each value within a column is unique. | | **Not null** | Validates that a column contains no null values. | | **Accepted values** | Validates that column values only include values from a defined set. | | **Relationships** | Validates referential integrity by ensuring that each value in a column exists as a corresponding value in another column or model. | ## How are test definitions defined? Test definitions use SQL queries to check your data. The query looks for problems in your data. If the query finds any rows (problems), the test fails. If the query returns no rows (no problems found), the test passes. ### Understanding test queries Think of a test query as a question you ask about your data: "Are there any rows that violate this rule?" If the answer is "yes" (rows are returned), the test fails. If the answer is "no" (no rows returned), the test passes. For example: * The `unique` test asks: "Are there any duplicate values in this column?" If duplicates exist, the query returns those duplicate rows and the test fails. * The `relationships` test asks: "Are there any values in this column that don't exist in the related table?" If mismatched values exist, the query returns those rows and the test fails. To create your own test definitions, **you need to know how to write SQL queries.** ### Example: Comparing row counts This example `equal_rowcount` test checks if two models have the same number of rows. It has the following parameters and definition: | Parameter | Type | | --------------- | ------- | | `model` | `table` | | `compare_model` | `table` | ```sql theme={null} with a as ( select count(*) as count_a from {{ model }} ), b as ( select count(*) as count_b from {{ compare_model }} ), final as ( select count_a, count_b, abs(count_a - count_b) as diff_count from a cross join b ) select * from final where diff_count > 0 ``` **How it works:** 1. The first CTE (`a`) counts rows in your model. 2. The second CTE (`b`) counts rows in the comparison model. 3. The `final` CTE calculates the difference between the two counts. 4. The final `SELECT` returns rows only when the difference is greater than 0. If the row counts match, the query returns no rows and the test passes. If they differ, the query returns a row showing the difference and the test fails. ## Build and run custom tests In the following sections, we'll build a test called `not_constant` to validate that a column does not have the same value in all rows. Follow the example to learn how to build a test definition and run the test. ### 1. Add a test definition To add a test definition, follow these steps: 1. In the left sidebar, click **+ Add Entity**. 2. Hover the **Tests** option and select **Test definitions**. 3. Assign a name to your test definition, such as `not_constant`. 4. Keep the default path `tests/generic`. This is the directory in the project Git repository where Prophecy will store the test definition. 5. Click **Create**. The test definition page opens. You can also create a new data test directly from the **Data Tests** tab of a table or model gem. ### 2. Define the test query On the test definition page, configure the following. 1. Under **Description**, add a summary of the test. 2. Under **Parameters**, add the parameters for the test. * All tests require the default `model` parameter of type `table`. * The `not_constant` test requires a `column_name` parameter of type `column` (the column to check). 3. Under **Definition**, add the SQL query for the test. ```sql theme={null} select count(distinct {{ column_name }}) as filler_column from {{ model }} having count(distinct {{ column_name }}) = 1 ``` Create a new model test definition ### 3. Assign the test definition After you've created a test definition, you can assign it to a table or model. 1. Open the table or model that you want to run the test on. 2. Click the **Data Tests** tab. 3. Click **+ New Test**. 4. Under **Data Test Type**, select the test definition you want to add to the gem. 5. Fill in the parameters for the test. Prophecy automatically sets the value of the `model` parameter to the current table. 6. Click **Create Test**. If necessary, you can add additional tests to the same table or model. If changes are made to the columns or schemas used in your data test, then Prophecy will delete the data test. For example, if you run into a data mismatch error on the Schema tab of your target model or update the schema, then your data test will be affected. If you want to change the failure condition of a test, you can do so by changing the **Advanced** settings for that test. | Setting | Description | | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Filter Condition** | When enabled, lets you define a filter expression for what you want your test to run on. Use the dropdown to select an expression. | | **Severity** | Determines whether the failure of the test returns an error or warning. Select from the dropdown to set the severity level. | | **Failure Calculation** | Expression that calculates a value from the test results. The **Error If** and **Warning If** conditions evaluate this calculated value. | | **Error If** | Condition that triggers an error. Evaluates the **Failure Calculation** result. If true and severity is `error`, the test fails with an error status. | | **Warn If** | Condition that triggers a warning. Evaluates the **Failure Calculation** result. If true, the test returns a warning status (regardless of severity setting). | | **Store Failures** | When enabled, stores all records that failed the test. The records are saved in a new table with schema `dbt_test__audit` in your database. The table is named after the name of the model and data test. Make sure you have write permission to create a new table in your data warehouse. | | **Set max no of failures** | When enabled, sets the maximum number of failures returned by a test query. You can set the limit to save resources and time by having the test stop its query as soon as it encounters a certain number of failed rows. | The Advanced settings are generally equivalent to the [Data Test gem parameters](/data-analysis/development/tests/project-tests#data-test-gem-parameters). ### 4. Run data tests Now that you've configured the tests, you can run them. 1. Click **Run all** to execute all the defined tests. Alternatively, run individual tests by clicking the play button on the test's row. 2. Once the tests finish running, you'll see the relevant test result next to each test definition. Assign a test definition to a table or model Click **View Log** to view the test logs. You'll only be able to view the log of the most recent test run. You can copy or download the logs if needed. ### 5. Schedule test runs When you schedule a pipeline, you can enable tests associated with Tables gems in the pipeline to run during that schedule. 1. Open a pipeline containing tests to schedule. 2. In the project header, click **... > Schedule**. 3. Edit the existing schedule or configure a new schedule. 4. Enable the **Run data quality tests** toggle. 5. Click **Confirm** to save the changes. Pipeline schedules differ from model schedules. Use the following steps if you are using **scheduling models only**. 1. In the left sidebar of the project, click **+ Add Entity**. 2. Click **Job**. This opens the **Create Job** dialog. 3. Enter a name for your job and click **Create New**. 4. Drag a **Model** gem to your visual canvas. 5. Click the model to open the model properties. 6. Select the database object you want to run the test on. 7. Select the **Run tests** checkbox in the left sidebar of the model gem. 8. Ensure that your **project, model**, and **fabric** are correct. 9. Click **Save**. ## Share test definitions If you publish your project as a [package](/data-analysis/development/extensibility/package-hub/package-hub), you can share your test definitions with other teams. Once someone imports the package, they will be able to use your test definitions in their own projects. We support importing test definitions (dbt generic tests) from dbt Hub packages, like `dbt_utils`. Learn how to import dbt packages in [Dependencies](/data-analysis/development/extensibility/dependencies). # What are data tests? Source: https://docs.prophecy.ai/data-analysis/development/tests/test-comparison Understand the purpose of different test types in Prophecy There are two types of data tests in Prophecy: table tests and project tests. Table tests and project tests serve different purposes in your data quality strategy. Understanding their differences helps you select the right approach for each validation scenario. The bottom line is that you can use **either type of test** to validate the same data quality requirements. However, table tests require parameterized SQL queries, whereas project tests leverage visual pipeline building. ## When to use table tests Use table tests when you need **reusable**, **parameterized** tests that can be applied across multiple tables or models. Table tests are written as "test definitions," which are essentially stored SQL queries with parameters. They are useful for: * **Parameterization**: Leverage parameters (like `model` and `column_name`). * **Sharing**: Share tests with other teams through Prophecy packages. * **Standardization**: Establish consistent data quality checks across your project. The test logic must be expressible as a **parameterized SQL query**. Because of this, you need to know how to write SQL queries. ## When to use project tests Use project tests when your validation logic doesn't need to be reused elsewhere. Project tests provide the following benefits: * **No SQL required**: Prepare data for testing using the visual pipeline builder. * **Workflow-specific**: Validate specific combinations of models or tables that are unique to a particular pipeline or workflow. However, project tests are not parameterized, so you cannot reuse the same test with different inputs. ## Example: Comparing row counts between two tables Both table tests and project tests can validate the same data quality requirement. This example demonstrates how to verify that two tables have the same number of rows using each approach. ### Scenario You need to ensure that `customers` and `customers_cleaned` tables have matching row counts. If the counts differ, the test should fail. ### Approach 1: Table test Create a parameterized test definition that you can apply to any table. 1. Create a new test definition. Equal row count table test 2. Define the parameters `{{ model }}` and `{{ compare_model }}` to accept any two tables. 3. Define the SQL query for the test. The query returns rows only when the counts differ (test fails). ```sql theme={null} with a as ( select count(*) as count_a from {{ model }} ), b as ( select count(*) as count_b from {{ compare_model }} ), final as ( select count_a, count_b, abs(count_a - count_b) as diff_count from a cross join b ) select * from final where diff_count > 0 ``` 4. Add the test to a table in a pipeline. * Prophecy sets the value of the `model` parameter to the current table. * You specify the `compare_model` parameter to the table you want to compare. Applying the table test to compare customers and customers_cleaned tables 5. Run the test. **Result:** You can reuse this test definition for any table in any pipeline in your project. ### Approach 2: Project test Build a visual pipeline that performs the same validation using gems instead of SQL. 1. Add two Table gems to the pipeline to read `customers` and `customers_cleaned`. 2. Add two Aggregate gems to count rows in each table. 3. Add a Join gem to combine the counts into a single row. 4. Add a Reformat gem to calculate the difference between counts. 5. Add a Filter gem to return rows only when the difference is greater than `0`. 6. Connect to the Data Test gem to evaluate the result: if any rows exist, the test fails. **Pipeline structure:** ``` Table 1 β†’ Aggregate (count rows) ┐ β†’ Join β†’ Reformat β†’ Filter β†’ Data Test Table 2 β†’ Aggregate (count rows) β”˜ ``` Project test pipeline showing visual gems for comparing row counts **Result:** This test validates the specific `customers` and `customers_cleaned` tables. To test different tables, you must create a new project test. # Unit tests for Data Analysis Source: https://docs.prophecy.ai/data-analysis/development/tests/unit-tests Create unit tests to validate the input and output of individual gems Private Preview Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Unit tests verify that a single component of your pipeline works correctly in isolation. When you create a unit test for a gem, you define specific input data and the expected output data or a set of predicates that must evaluate to true for the test to pass. Think of unit tests as a way to document and verify the behavior of each transformation step. For example, if you have a Join gem that combines customer and order data, you can create a unit test that: * Provides sample customer rows and order rows as input * Defines the expected joined output rows * Verifies that the join condition works correctly When you modify the Join gem's configuration later, running the unit test confirms that your changes didn't break the expected behavior. Unit tests catch errors early, before they affect downstream pipelines or production data. ## Prerequisites Unit tests are available only for [Simplified PySpark](/data-analysis/development/projects/project-languages#simplified-pyspark) projects. ## Types of unit tests You can configure two types of unit tests on gems. * [Output rows equality](#output-rows-equality): Compares the actual output rows against a saved snapshot of expected data. * [Output predicates](#output-predicates): Evaluates Spark expressions against output data to verify business rules and constraints. ### Output rows equality Output rows equality tests compare the actual output rows against expected data you define. Use this test type when you need to verify that transformations produce identical results. 1. Open the gem you want to test. 2. Click **Unit Tests** in the gem configuration. 3. Click **Create Test** to add a new unit test. 4. In the **Settings** section, select **Output rows equality** from the dropdown. 5. Click one or more columns in the left panel to add them to the **Selected Columns** table. 6. Click **Create**. 7. Define expected input data: * Select an input port tab, such as `in0` or `in1`. * Select the correct data type for each column you are testing. * Click **+ Add Row** to add expected input rows. * Enter values for each column. 8. Define expected output data: * Select the `out` port. * Select the correct data type for each column you are testing. * Click **+ Add Row** to add expected output rows. * Enter values for each column. 9. Click **Done** to save the unit test. ### Output predicates Output predicates let you define expressions that must evaluate to true for the test to pass. Use predicates when you need to validate business rules, data constraints, or complex conditions rather than exact row matches. 1. Open the gem you want to test. 2. Click **Unit Tests** in the gem configuration. 3. Click **Create Test** to add a new unit test. 4. In the **Settings** section, select **Output predicates** from the dropdown. 5. Click one or more columns in the left panel to add them to the **Selected Columns** table. 6. Click **Create**. 7. Define expected input data: * Select an input port tab, such as `in0` or `in1`. * Select the correct data type for each column you are testing. * Click **+ Add Row** to add expected input rows. * Enter values for each column. 8. Add predicates for the output: * In the predicates table, enter a **Predicate Name** in the first column. Use descriptive names that indicate what the predicate validates. * Enter an expression in the **Expression** column. The expression must evaluate to a boolean value and return `true` for the test to pass. * Click in an empty row below to add additional predicates if needed. 9. Click **Done** to save the unit test. #### Example predicates Review the following example predicates to help you understand how to write predicates. | Predicate Name | Expression | | --------------------------------- | ----------------------------------------------------------- | | Amount is positive | `amount > 0` | | First name differs from last name | `first_name != last_name` | | Order date in valid range | `order_date >= '2024-01-01' AND order_date <= '2024-12-31'` | You can add multiple predicates to a single unit test. All predicates must evaluate to `true` for the test to pass. ## Generate sample data automatically Enable automatic data generation to create test input data without manually entering rows. This option generates sample rows from upstream data. 1. In the unit test configuration, toggle on **Generate Data**. 2. Enter the number of rows to generate in the **Rows** field. Prophecy samples this many rows from the input data. 3. Click **Create** to generate the sample input data. 4. Review the generated sample input data. Edit the data if needed. 5. Click **Done** to save the unit test. # Clone projects Source: https://docs.prophecy.ai/data-analysis/development/versioning/clone-projects Create independent copies of existing projects Clone a project to create a copy of that project. Cloning creates a completely independent project with: * A new Git repository containing a copy of the project code * Separate version history going forward * Independent development that won't affect the original project Unlike [importing](/data-analysis/development/versioning/import-projects), which connects to an existing repository, cloning creates a fresh repository for the new project. ## Create a clone To create a clone: 1. Navigate to **Metadata > Projects**. 2. Open the project you want to clone. 3. Click **... > Clone**. 4. Configure the clone settings: * **Name**: Display name for the cloned project * **Team**: The team that will own the clone * **Git Storage Model**: Choose between Simple or Normal [Git model](/data-analysis/development/versioning/version-control#version-control-options) * **Git account**: Select which Git account will host the new repository * **Repository details**: If using external Git, specify where the new repository will be created * **Copy all release tags** (optional): Check this to transfer release versions to the clone 5. Click **Clone Project**. You will see the new project appear in your chosen team. Prophecy only copies the main branch from the source project when cloning. If you have uncommitted changes or changes on other branches, publish the project before cloning (merge changes into main). Otherwise, those changes won't appear in the cloned project. # Version conflicts Source: https://docs.prophecy.ai/data-analysis/development/versioning/conflicts How to prevent and resolve version conflicts from multiple contributors When you work on the same project as someone else, you're working on the same copy of the project. To prevent multiple users from overwriting each other's work, Prophecy limits editing of any entity (pipeline, table, document, etc.) to one user at a time. ## What happens when two users are working on the same project simultaneously? When other users are active, you can see their presence indicated by a user icon on each entity. You can view these entities, but you can't make changes until the other user closes the entity or you request to take over. Active users in the project To take over editing control from another user: 1. Open the entity you want to edit. 2. Click **Take Over** in the canvas footer. 3. Wait for the other user to respond. If the other user is idle and doesn't respond within 30 seconds, Prophecy grants you control automatically. ## How does this relate to version control? Prophecy projects are versioned using Git. We've simplified Git's typical branching workflow into a straightforward save and publish workflow. As a result, everyone works on the same copy of a project, which is why simultaneous editing isn't allowed. In a typical Git workflow, every user gets their own copy of a project. When individual work is done, users merge their copies into one place. If changes overlap, users must resolve the conflicts manually. Prophecy's simplified workflow avoids conflicts entirely by preventing simultaneous edits. To learn more about the typical Git workflow and how it integrates with Prophecy, see [Git](/data-engineering/ci-cd/git/git). ## I see a conflict. How did this happen? Conflicts should not occur on projects using Simple Git; however, they can happen if you are using an external Git repository. This means there is another place where changes can be made to the project, and Prophecy can't prevent them. (Typically, Simple Git projects are hosted on Prophecy-managed Git, so this is not a concern. Users can't edit the project code directly in the Prophecy-managed Git repository.) To be more specific, if a user makes a conflicting change to the `dev` or `main` branch of the external Git repository: * You will have to pull (integrate) these changes from the external Git into your project in Prophecy. * These external changes may conflict with changes made in your project in Prophecy, and you will need to resolve them. To resolve conflicts, you can [use the Prophecy interface](/data-engineering/ci-cd/git/git-resolve) to resolve them or resolve them in your external Git repository. # Git storage models Source: https://docs.prophecy.ai/data-analysis/development/versioning/git-storage-model Understand how different Git models work in Prophecy Normal and Fork Git models are only available on the [Express and Enterprise Editions](/data-analysis/administration/platform/editions). By default, new projects use the Simple Git Storage Model and are hosted on Prophecy-managed Git. This lets users develop projects without knowledge of Git or branching strategies. While this is sufficient for most users, advanced users may want to change the Git Storage Model during project creation to use a normal or forked Git repository. Git setup during project creation ## Model comparison The following table describes the different Git storage models and how they work. | Git Storage Model | Description | | ----------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | Simple | Provides an intuitive visual workflow for project drafting and publication. Users all work on the same `dev` branch in the Git backend. | | Normal (no forks) | Enables the typical Git workflow aligned with DevOps best practices. Users all work in the same repository on different branches. | | Fork per user | (External Git only) Enables the typical Git workflow aligned with DevOps best practices. Users work on their own copy of the repository. | Regardless of the Git storage model you choose, you'll be able to use a Prophecy-managed Git repository or your own external Git repository to host your project code. ## Simple Git Storage Model As you move through the Simple versioning workflow in your project, Prophecy actually maps these actions to Git processes in the backend. In other words, actions like saving, publishing, and restoring changes trigger Git commands. This is possible because all Prophecy projects are hosted on Git, regardless of the project's Git storage model. The following diagram explains what each versioning action does in Git. If you connect to an external Git provider (rather than use Prophecy-managed Git), you can view how each action in is reflected in Git as you work on your project. Simple Git The table below reiterates the diagram. | Action in Prophecy | Action in Git | | ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Save to draft |
  • Pull changes from the remote `dev` branch
  • Pull changes from the local `main` branch
  • Commit changes to local `dev`
  • Push changes to remote `dev`
| | Publish |
  • Merge changes into local `main`
  • Add a Git Tag with the published version number
  • Push changes to remote `main`
| | Restore previous version |
  • Run `git reset --soft`
  • Commit the changes to revert in `dev`
  • Push changes to remote `dev` branch
(You must manually publish the project again to merge what you reverted to `main`) | # Import projects Source: https://docs.prophecy.ai/data-analysis/development/versioning/import-projects Import projects from existing Git repositories Import existing Prophecy projects from external Git repositories into your workspace. This connects your Prophecy project to the repository without creating a copy of the repository. You can only import projects from external Git providers (GitHub, GitLab, Bitbucket, etc.). Projects stored in Prophecy-managed Git cannot be imported using this method. ## Import a project To import a project from Git: 1. Open **Create Entity** from the left sidebar. 2. Hover over the **Project** tile. 3. Click **Import**. 4. Connect to the Git account that hosts your project repository. 5. Select your repository from the **Repository** dropdown. If your repository doesn't appear (common with forked repositories), paste the repository URL directly into the Repository field. 6. Verify the **Default Branch** field auto-populates correctly. 7. Keep the default root path or specify a path to the project folder within the repository. 8. Click **Continue**. 9. Configure your project settings: * **Name**: Display name for the project in Prophecy * **Description**: Optional context about the project * **Team**: The team that will own this project * **Project type**: Prophecy automatically sets the language based on the repository content (cannot be changed) * **Provider**: Where pipelines will execute 10. Click **Complete**. ## Import vs. clone Importing a project connects directly to the existing repository without creating a new copy. This means: * Multiple Prophecy projects can point to the same Git repository * Changes made in any connected project affect the shared repository * All projects connected to the same repository will see each other's commits Use caution when multiple projects share a repository, as changes are not isolated between projects. # Migrate project repository from Prophecy-managed Git to external Git Source: https://docs.prophecy.ai/data-analysis/development/versioning/migrate-managed Move the repository of a project from managed-Git to external Git This guide explains how to migrate a project that was started in Prophecy-managed Git to an external Git repository (such as Bitbucket or GitHub). You'll get an understanding of the process, the reasons behind this workflow, and the prerequisites for a smooth migration. This process might be helpful when you want to: * Allow analysts to quickly create and iterate on projects in Prophecy, even if they lack credentials for the company's external Git. * Allow the engineering team to manage CI/CD and code promotions as per company standards in an external repository once pipelines are ready for production. Migrating a project's repository is not the same as cloning a project. A cloned project is a new, distinct project with its own repository. Migration moves the existing project and its ID to the new repo. ## Prerequisites Before migrating a project repository from Prophecy-managed Git to an external Git provider, ensure the following: * The original project must be a SQL project that uses the [Simple Git Storage Model](/data-analysis/development/versioning/version-control#version-control-options). * You must be either the project owner (the user who created the project) or a team admin of the team associated with the project. * You must have valid credentials for the external Git provider. If you don't, team members can [share credentials](/data-engineering/ci-cd/git/git#share-credentials) with you in Prophecy. * You must have access to an empty repository in the external Git provider where the project will be migrated. * The external repository must contain a dedicated branch that will receive updates when you publish the project from Prophecy. (This branch doesn't have to be named `main`; you can choose any branch to serve as the main publishing branch.) ## Step-by-step guide ### Initiate the migration Once you've confirmed all prerequisites are in place, you can begin the migration process from within your Prophecy project. 1. Open your Prophecy project. 2. Click the `...` menu in the top right corner. 3. Select **Set up on External Git**. 4. Select the credential you wish to use. ### Fill migration details After initiating the migration, you'll be prompted to provide the following information about the external repository. 1. Fill in the following fields: * **Repository** β€” Select the repository from the external Git provider to store the project code. * **Path** β€” Write a non-root path where the project code will be stored (for example, `/myproject`, not just `/`). * **Main Branch** β€” Reference an existing branch in the external repo that will be the project publication branch. * **Development Branch** β€” Name the branch where users will continue to do development work in the project. * **Copy all release tags** - Option to migrate release tags to the new repo. 2. Click **Set up on External Git** to confirm the migration. Prophecy will move the entire project's Git history to the path (directory) defined in the external repo in the specified development branch. ### Post-migration After the migration, your team can continue collaborating on the project using shared credentials in Prophecy, ensuring that analysts and engineers maintain access without requiring individual Git credentials. With the project now hosted in your external Git provider, you can integrate it into your organization's CI/CD workflows for code promotion, security scanning, and managed deployment. # Version control Source: https://docs.prophecy.ai/data-analysis/development/versioning/version-control Save and view project history Versioning in Prophecy helps teams track changes, collaborate efficiently, and roll back when needed. It also supports auditing and compliance by keeping a clear, versioned history of all updates. This page details the stages of the visual workflow. Simple version menu ## Workflow The following sections describe the versioning workflow, where Prophecy creates a linear version history per project where you can audit changes, see collaborator activity, and revert to previous versions. ### Save to draft As you develop your project, Prophecy **automatically** preserves your changes. However, we recommend periodically saving your changes as drafts. To do so: 1. Open the version menu in the Studio header. 2. Click **Save to draft**. 3. Fill out the **Version description** to summarize the changes made since the last saved version, or let AI generate one for you. 4. Review your changes in the **Changes since last saved** section. 5. Click **Save**. Your changes are now saved as a draft. ### Publish a new version When you're ready to use your project in production, you'll publish it. For an in-depth review of the publication process, see [project publication](/data-analysis/production/publication). ### Show version history Prophecy tracks different versions of your project that you save and publish. To access the version history: 1. Open the version menu in the Studio header. 2. Click **Show version history**. The top level versions represent the published versions of the project. Expand each version to see the drafts that were saved for that version. Each version displays the author and time since the version was saved. ### Restore previous version To restore a previous version: 1. Open the version menu in the Studio header. 2. Click **Show version history**. 3. Find the version you want to restore. (This can be a published version or a draft.) 4. Hover the version and click **... > Restore this version**. ### Publish a previous version To publish a previous version: 1. Open the version menu in the Studio header. 2. Click **Show version history**. 3. Find the version you want to publish. (This can be a published version onlyβ€”not a draft.) 4. Hover the version and click **... > Publish**. ## What's next Learn about how Prophecy uses Git to version your projects in [Git storage models](/data-analysis/development/versioning/git-storage-model). # Azure Data Lake Storage Source: https://docs.prophecy.ai/data-analysis/environment/connections/adls Connect to Azure Data Lake Storage (ADLS) accounts and containers Prophecy supports direct integration with [Azure Data Lake Storage](https://learn.microsoft.com/en-us/azure/storage/blobs/data-lake-storage-introduction) (ADLS), allowing you to read from and write to ADLS containers as part of your data pipelines. This page explains how to configure the connection, what permissions are required, and how ADLS connections are managed and shared within your team. ## Prerequisites Prophecy connects to ADLS using the credentials you provide. These are used to authenticate requests and authorize all file operations during pipeline execution. To ensure Prophecy can read from and write to your storage account, you must have the following **Azure RBAC role** or equivalent permissions: * **Storage Blob Data Contributor**: Read, write, and delete access to Blob storage containers and blobs. To learn more, see [Access control model in Azure Data Lake Storage](https://learn.microsoft.com/en-us/azure/storage/blobs/data-lake-storage-access-control-model). ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | -------------------------------------------------------------------------------------------------------------------------------------- | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/file/adls) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/file/adls) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Trigger scheduled pipeline upon [file arrival or change](/data-analysis/production/scheduling/triggers#file-arrival-or-change-trigger) | Yes | ## Connection parameters To create a connection with your ADLS account, enter the following parameters: | Parameter | Description | | ----------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | | Connection Name | Unique name for the connection. | | Client ID | Your Microsoft Entra app client ID | | Tenant ID | Your Microsoft Entra [tenant ID](https://learn.microsoft.com/en-us/entra/fundamentals/how-to-find-tenant) | | Client Secret ([Secret required](/data-analysis/environment/secrets/secrets)) | Your Microsoft Entra app client secret | | Account Name | Name of your ADLS [storage account](https://learn.microsoft.com/en-us/azure/storage/blobs/create-data-lake-storage-account) that hosts the container | | Container Name | Name of the [container](https://learn.microsoft.com/en-us/azure/storage/blobs/storage-blobs-introduction#containers) within the storage account | # Google BigQuery Source: https://docs.prophecy.ai/data-analysis/environment/connections/bigquery Learn how to connect to BigQuery A BigQuery connection allows Prophecy to access tables in your BigQuery project. When configured as a SQL Warehouse connection, BigQuery also provides the compute engine used to run pipeline transformations. This page explains how to use and configure a Google BigQuery connection in Prophecy. ## Prerequisites Prophecy connects to BigQuery using the credentials you provide. These credentials are used to authenticate your session and authorize all data operations during pipeline execution, including reading from and writing to tables. To use a BigQuery connection effectively, your user or service account should have: * `OWNER` dataset role to be able to read, insert, update, and delete datasets. To learn more, visit [Basic roles and permissions](https://cloud.google.com/bigquery/docs/access-control-basic-roles) in the BigQuery documentation. ## Connection type Prophecy supports BigQuery in two different roles. | Connection type | Description | | ----------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **SQL Warehouse connection** | BigQuery acts as the compute engine for a fabric. Prophecy generates SQL and BigQuery executes the transformations. To learn more about SQL Warehouse connections, see [Prophecy fabrics](/data-analysis/environment/fabrics/prophecy-fabrics). | | **Ingress/Egress connection** | BigQuery is used only as a data source or target. Prophecy reads data from or writes data to BigQuery, but pipeline transformations run in another warehouse. | ## When BigQuery is a fabric ``` Pipeline ↓ Fabric ↓ BigQuery executes SQL ↓ Tables updated in BigQuery ``` ## When BigQuery is ingress/egress ``` Pipeline ↓ Warehouse executes SQL (Databricks / Prophecy / Snowflake) ↓ Prophecy Automate ↓ BigQuery tables ``` ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | SQL Warehouse | Ingress/Egress | | ------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------- | -------------- | | Run SQL queries | Yes | No | | [Read](/data-analysis/gems/source-target/table/bigquery-read) and [write](/data-analysis/gems/source-target/table/bigquery-write) data with a Table gem | Yes | No | | Read data with a [Source gem](/data-analysis/gems/source-target/external-table/bigquery) | Yes | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/external-table/bigquery) | Yes | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | Yes | | Index tables in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | Yes | ## Connection parameters To create a connection with BigQuery, enter the following parameters. | Parameter | Description | | --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Connection Name | A unique name to identify the connection. | | Project ID | The ID of your [Google Cloud project](https://cloud.google.com/resource-manager/docs/creating-managing-projects). | | Dataset | The default location for target tables and [temporary tables](/data-analysis/development/runs/execution#external-data-handling).
Requires write permissions. | | Authentication Method | The method used to authenticate with BigQuery.
See [Authentication methods](#authentication-methods) for details. | | Bucket Name | A [Google Cloud Storage bucket](https://cloud.google.com/storage/docs/buckets) used for write optimization (recommended).
When specified, Prophecy writes data to the bucket, then loads it into BigQuery.
Note: Loading data from a bucket offers better performance than writing with the [BigQuery API](https://cloud.google.com/bigquery/docs/reference/rest) (default). | ## Authentication methods You can authenticate your BigQuery connection using either [OAuth](#oauth-user-to-machine) or a [Private Key](#private-key-machine-to-machine). Each method grants Prophecy the ability to read and write data in your BigQuery environment based on the [Identity and Access Management (IAM)](https://cloud.google.com/iam/docs/overview) roles assigned to the authenticated identity. ### OAuth (User-to-Machine) OAuth is a user-based authentication (U2M) method best suited for interactive pipeline development. It allows each user to sign in with their own Google account, which ensures that data access is governed by their individual IAM roles and permissions. To leverage user-based OAuth: 1. Under **Authentication method**, select **OAuth**. 2. Under **App Registration**, select the correct app registration or use the default. If no app registrations appear, an admin must configure an [OAuth app registration](/data-analysis/administration/management/cluster-admin-settings/oauth-setup). When your connection is configured to use OAuth, the following occurs when a user attaches the fabric to their project: 1. The user is prompted to sign in with their Google account. 2. Prophecy uses the user's credentials to authenticate the connection. 3. The connection operates with the user's IAM roles and permissions. 4. Token management, including refresh, is handled automatically by Google. The default [refresh token expiration](https://developers.google.com/identity/protocols/oauth2#expiration) time is 7 days. For more about OAuth and how it works with Google Cloud, see [Using OAuth 2.0 to Access Google APIs](https://developers.google.com/identity/protocols/oauth2). ### Private Key (Machine-to-Machine) Use a Service Account when you want a non-user identity for authentication (M2M). This is ideal for automated or shared processes that require stable, long-term access without re-authentication interruptions. 1. Create and download a [Service Account Key](https://developers.google.com/workspace/guides/create-credentials#service-account) from the Google Cloud console. 2. Paste the full JSON content into a [Prophecy secret](/data-analysis/environment/secrets/secrets) as text. Binary upload is not supported. 3. Open a BigQuery connection. 4. Under **Authentication method**, select **Private Key**. 5. Use the Prophecy secret in the **Service Account Key** field. This method allows all team members with access to the fabric to use the connection in their projects. Those users inherit the access and permissions of the Service Account, as defined in its IAM roles. ## Data type mapping When Prophecy processes data from Google BigQuery using an external SQL warehouse, it converts BigQuery data types to a compatible type. | BigQuery | Databricks | | ---------- | ----------------------------------- | | STRING | STRING
Alias: String | | BYTES | BINARY
Alias: Binary | | NUMERIC | DECIMAL128(38,9)
Alias: Bigint | | BIGNUMERIC | DECIMAL256(38,9)
Alias: Bigint | | FLOAT | DOUBLE
Alias: Double | | RECORD | STRUCT
Alias: Struct | | ARRAY | ARRAY
Alias: Array | | INTEGER | INT
Alias: Integer | | BOOLEAN | BOOLEAN
Alias: Boolean | | DATE | DATE
Alias: Date | | DATETIME | TIMESTAMP
Alias: Timestamp | | TIME | STRING
Alias: String | | GEOGRAPHY | STRING
Alias: String | | INTERVAL | STRING
Alias: String | | JSON | STRING
Alias: String | Learn more in [Supported data types](/data-analysis/gems/data-types). # Ingest and write data with connections Source: https://docs.prophecy.ai/data-analysis/environment/connections/connections Use connections to read and write data from external sources Prophecy lets you work with various data providers when building your pipelines. To read and write data from external sources, create **connections** inside a [Prophecy fabric](/data-analysis/environment/fabrics/prophecy-fabrics). When you attach to a fabric with connections, you can: * Reuse credentials that are established in the connection. * Browse data from the data provider in the [Environment browser](/data-analysis/development/studio/studio#sidebar) of your Prophecy project. Most connections are only used to read from and write to data sources. The SQL Warehouse connection is an exception: it also provides the compute environment for pipeline execution. ## Connections access Prophecy controls access to connections through fabric-level permissions. To access a connection, you must have access to the fabric that contains the connection. You can only access fabrics that are assigned to one of your teams. ```mermaid theme={null} erDiagram Project }o--o{ Fabric : uses Team ||--o{ Project : owns Team ||--o{ Fabric : owns ``` In this diagram, the lines show relationships between entities. `||--||` indicates a one-to-one relationship, `||--o{` indicates a one-to-many relationship (one entity on the left can relate to many entities on the right), and `}o--o{` indicates a many-to-many relationship. ## Add a new connection To configure a new connection in a Prophecy fabric: 1. Open the **Metadata** page from the left sidebar in Prophecy. 2. Navigate to the **Fabric** tab. 3. Open the fabric where you want to add the connection. 4. Navigate to the **Connections** tab. 5. Click **+ Add Connection**. This opens the **Create Connection** dialog. 6. Select a data provider from the list of connection types. 7. Click **Next** to open the connection details. 8. Configure the connection and save your changes. Learn about individual connection parameters in the connection's respective reference page. To make connections configurable across projects or pipelines, use [connection parameters](/data-analysis/development/parameters/parameters). ## View connection status Connection status indicates whether Prophecy can reach your data sources through network connectivity tests. Status appears in the **Connections** tab of your fabric. Each connection displays one of four statuses: * **Connected**: Network connectivity test passed. You can use the connection. * **Disconnected**: Network connectivity test failed. Common causes include a firewall blocking the connection or expired credentials. * **Inactive**: The connection hasn't been used within the past week. Prophecy suspends network connectivity testing for inactive connections to reduce overhead. Testing resumes automatically when you use the connection again. * **Not tested**: This indicates that Prophecy does not support network connectivity testing for this connection type. ## What's next Visit the following pages for details about individual connections. # Databricks Source: https://docs.prophecy.ai/data-analysis/environment/connections/databricks Learn how to connect with Databricks This page explains how to use and configure a Databricks connection in Prophecy. A Databricks connection allows Prophecy to access files, tables, and compute resources in your Databricks workspace. You can use the same connection to access any of these resources, as long as the authenticated account in the connection has the appropriate permissions. ## Prerequisites Prophecy connects to Databricks using the credentials you provide. These credentials are used to authenticate your session and authorize all data operations during pipeline execution, including reading from and writing to tables. To use a Databricks connection effectively, your user or service principal must have the following: * [Basic table permissions](https://docs.databricks.com/aws/en/tables/#basic-table-permissions) defined in the Databricks documentation. * Additional [Unity Catalog privileges](https://docs.databricks.com/aws/en/data-governance/unity-catalog/manage-privileges/privileges): * `CREATE VOLUME` for permission to create the `PROPHECY_ORCHESTRATOR_VOLUME` * `READ VOLUME` for the path `/Volumes///PROPHECY_ORCHESTRATOR_VOLUME` * `WRITE VOLUME` for permission to delete intermediate files from the volume ## Connection type Prophecy supports Databricks as both a SQL Warehouse connection and an Ingress/Egress connection. To learn more about these different connection types, visit [Prophecy fabrics](/data-analysis/environment/fabrics/prophecy-fabrics). ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------- | | Run SQL queries | Yes β€” SQL Warehouse Connection only | | [Read](/data-analysis/gems/source-target/table/databricks-read) and [write](/data-analysis/gems/source-target/table/databricks-write) data with a Table gem. | Yes β€” SQL Warehouse Connection only | | Read data with a [Source gem](/data-analysis/gems/source-target/external-table/databricks) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/external-table/databricks) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Index tables in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | Yes | ## Connection parameters To create a connection with Databricks, enter the following parameters. | Parameter | Description | | ------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Connection Name | Name to identify your connection. | | JDBC URL | URL to connect to your SQL warehouse. For example: `jdbc:databricks://:443/default;transportMode=http;ssl=1;AuthMech=3;httpPath=/sql/1.0/warehouses/` | | Catalog | Default write location for target tables. | | Schema | Default write location for target tables. | | Authentication method | How you want to authenticate your Databricks account. Learn more in [Authentication methods](#authentication-methods). | | Select a cluster to run the script on | Cluster where [Script gems](/data-analysis/gems/custom/script) run. Specify a cluster if you want to install libraries on your cluster to use in your scripts.
The dropdown will only display running clusters (inactive clusters will not appear).
If not specified, Script gems run on Databricks Serverless. | When you use Databricks as your primary SQL warehouse, Prophecy also uses the catalog and schema you define in the connection to store temporary tables during [pipeline execution](/data-analysis/development/runs/execution#external-data-handling). Therefore, you must have write access to the schema in Databricks. To avoid conflicts, define distinct catalog and schema locations for each fabric. ## Authentication methods You can configure your Databricks connection using the following authentication methods. ### OAuth When you select OAuth as your authentication method for the connection, Prophecy can authenticate using either user-based (U2M) or service principal-based (M2M) OAuth. A single fabric cannot use both methods for the same connection. To use both U2M and M2M, create separate fabrics. | OAuth Type | Authentication | Requirements | Token Expiry | Best For | | --------------------------- | ----------------------------- | ----------------------------------------------------------------------------------------------- | ----------------------------------- | -------------------------- | | **User-based (U2M)** | Individual user accounts | [App Registration](/data-analysis/administration/management/cluster-admin-settings/oauth-setup) | Periodic re-authentication required | Interactive development | | **Service Principal (M2M)** | Service principal credentials | Service Principal Client ID and Service Principal Client Secret | No expiration | Scheduled jobs, automation | You can schedule pipelines with user-based OAuth, but you'll need to re-authenticate periodically. Prophecy estimates token expiry based on Databricks' response and prompts you when re-authentication is needed. To avoid interruptions, switch to service principal-based OAuth for production workloads. Use different fabrics for development and production to align authentication with your environment's needs. | Environment | OAuth Type | Description | | --------------- | ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Development** | User-based (U2M) | Developers log in with personal Databricks accounts. Reflects individual permissions and enhances security through regular re-authentication. | | **Production** | Service Principal (M2M) | Pipelines run under a shared service principal for continuous, unattended operation. Access should be restricted to trusted users since credentials are shared. | ### Personal Access Token (PAT) When you choose **Personal Access Token** (PAT) for the authentication method, you'll authenticate using a [Databricks personal access token](https://docs.databricks.com/aws/en/dev-tools/auth/pat). When you set up the connection, you will use a [secret](/data-analysis/environment/secrets/secrets) to enter your PAT. Using the PAT authentication method: * All team members who have access to the fabric can use the connection in their projects. * No additional authentication is required. Team members automatically inherit the access and permissions of the stored connection credentials. ## Data type mapping When Prophecy processes data from Databricks using an external SQL warehouse, it converts Databricks data types to compatible types. | Databricks | BigQuery | | ---------- | ------------------------------- | | INT | INT64
Alias: Integer | | TINYINT | INT64
Alias: Integer | | SMALLINT | INT64
Alias: Integer | | BIGINT | INT64
Alias: Integer | | STRING | STRING
Alias: String | | BOOLEAN | BOOL
Alias: Boolean | | DECIMAL | NUMERIC
Alias: Numeric | | FLOAT | FLOAT64
Alias: Float | | DOUBLE | FLOAT64
Alias: Float | | BINARY | BYTES
Alias: Bytes | | TIMESTAMP | TIMESTAMP
Alias: Timestamp | | DATE | DATE
Alias: Date | | MAP | JSON
Alias: JSON | | ARRAY | ARRAY
Alias: Array | | STRUCT | STRUCT
Alias: Struct | | VOID | BYTES
Alias: Bytes | | VARIANT | STRUCT
Alias: Struct | Learn more in [Supported data types](/data-analysis/gems/data-types). # Google Cloud Storage Source: https://docs.prophecy.ai/data-analysis/environment/connections/gcs Learn how to connect to Google Cloud Storage (GCS) buckets Prophecy supports direct integration with Google Cloud Storage (GCS), allowing you to read from and write to GCS buckets as part of your data pipelines. This page explains how to configure the connection, what permissions are required, and how GCS connections are managed and shared within your team. ## Prerequisites Prophecy connects to GCS using a Google Cloud service account key that you provide. This key is used to authenticate requests and authorize all file operations during pipeline execution. To ensure Prophecy can read from and write to GCS as needed, the service account must have the following permissions: * `storage.objects.list` β€” to list the contents of the bucket * `storage.objects.get` β€” to read files from the bucket * `storage.objects.create` β€” to write files to the bucket To learn more, visit [IAM permissions for Cloud Storage](https://cloud.google.com/storage/docs/access-control/iam-permissions) in the Google Cloud documentation. ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | -------------------------------------------------------------------------------------------------------------------------------------- | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/file/gcs) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/file/gcs) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Trigger scheduled pipeline upon [file arrival or change](/data-analysis/production/scheduling/triggers#file-arrival-or-change-trigger) | Yes | ## Connection parameters To create a connection with your GCS buckets, enter the following parameters: | Parameter | Description | | ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Connection Name | Unique name for the connection. | | Service Account Key ([Secret required](/data-analysis/environment/secrets/secrets)) | Key used to authenticate the connection.
See [Create and delete service account keys](https://cloud.google.com/iam/docs/keys-create-delete) for more information. | | Project ID | Google Cloud project ID that owns the bucket. | | Bucket Name | Name of your GCS bucket. | **Service Account Key** Paste the full JSON content of your GCP service account key into a [Prophecy secret](/data-analysis/environment/secrets/secrets) as text. Binary upload is not supported. ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once a GCS connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored service account credentials. Be mindful of the access level granted by the stored service account key. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # SAP HANA Source: https://docs.prophecy.ai/data-analysis/environment/connections/hana Learn how to connect to SAP HANA Prophecy supports direct integration with SAP HANA, allowing you to read from and write to the database as part of your data pipelines. This page explains how to configure the connection, what permissions are required, and how Hana connections are managed and shared within your team. ## Prerequisites To connect Prophecy to SAP HANA, you need: * **Database credentials**: Prophecy uses the credentials you provide to authenticate your session and authorize all read and write operations during pipeline execution. * **User permissions**: Your SAP HANA account must have the required [object privileges](https://learning.sap.com/learning-journeys/installing-and-administering-sap-hana/describing-sap-hana-privileges-and-roles) to read from and write to the database. * **Network setup**: A PrivateLink connection between the Prophecy network and your SAP HANA network. Contact Prophecy support for help setting up PrivateLink. ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/external-table/hana/hana) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/external-table/hana/hana) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | ## Connection parameters To create a connection to SAP HANA enter the following parameters. | Parameter | Description | | --------------------- | ------------------------------------------------- | | Connection Name | A unique name for the connection. | | Authentication Method | Select **Username and Password** or **User Key**. | ### Authentication methods #### Username and Password | Parameter | Description | | --------- | -------------------------------------------------- | | Host | The IP address or hostname of the SAP HANA server. | | Port | The port number used to connect to the server. | | Username | Your SAP HANA username. | | Password | Your SAP HANA password. | #### User Key | Parameter | Description | | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | User Key | The HDB user store key to authenticate the connection. See [SAP HANA User Store (hdbuserstore)](https://help.sap.com/docs/SAP_HANA_CLIENT/f1b440ded6144a54ada97ff95dac7adf/708e5fe0e44a4764a1b6b5ea549b88f4.html?version=2.24) for details. | ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once a Hana connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # MongoDB Source: https://docs.prophecy.ai/data-analysis/environment/connections/mongodb Learn how to connect with MongoDB MongoDB is a NoSQL database designed to store and retrieve unstructured or semi-structured data using BSON documents. ## Prerequisites When you create a MongoDB connection in Prophecy, access permissions are tied to the credentials you use. This is because Prophecy uses your credentials to execute all data operations, such as reading from or writing to collections. To fully leverage a MongoDB connection in Prophecy, you need the following MongoDB permissions: * `Read` from the collection defined in the connection * `Write` to the collection defined in the connection ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/external-table/mongodb) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/external-table/mongodb) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Index tables in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | ## Connection parameters To create a connection with MongoDB, enter the following parameters: | Parameter | Description | | ------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------- | | Connection Name | Name to identify your connection | | Protocol | Protocol to use to communicate to the database
Example:`mongodb+srv` for cloud-hosted clusters | | Host | Where your MongoDB instance runs
Example:`cluster0..mongodb.net` for cloud-hosted clusters | | Username | Username for your MongoDB instance | | Password ([Secret required](/data-analysis/environment/secrets/secrets)) | Password for your MongoDB instance | | Database | Default database for reading and writing data | | Collection | Collection to use for the connection | ## Data type mapping When Prophecy processes data from MongoDB using SQL warehouses, it converts MongoDB-specific data types to formats compatible with your target warehouse. This table shows how [MongoDB data types](https://www.mongodb.com/docs/manual/reference/bson-types/) are transformed for Databricks, BigQuery, and Snowflake. | MongoDB | Databricks | BigQuery | Snowflake | | ------------------------------------------------------------------------------------------- | ----------------------------------- | -------------------------------------- | --------------------------------- | | 32-bit integer | `int`
Alias: Integer | `int64`
Alias: Integer | `number`
Alias: Number | | 64-bit integer | `bigint`
Alias: Bigint | `int64`
Alias: Integer | `number`
Alias: Number | | Double | `double`
Alias: Double | `float64`
Alias: Float | `float`
Alias: Float | | Boolean | `boolean`
Alias: Boolean | `bool`
Alias: Boolean | `boolean`
Alias: Boolean | | String | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | Date | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | Binary | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `binary`
Alias: Binary | | ObjectId | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | Decimal128 | `decimal(34,2)`
Alias: Decimal | `bignumeric(34,2)`
Alias: Numeric | `number(34,2)`
Alias: Number | | Min key | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | Max key | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | Regular Expression | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | DBPointer | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | JavaScript | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | JavaScript with scope | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | Symbol | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | [Embedded Document](https://www.mongodb.com/docs/manual/tutorial/query-embedded-documents/) | `struct`
Alias: Struct | `struct`
Alias: Struct | `variant`
Alias: Variant | | Array | `array`
Alias: Array | `array`
Alias: Array | `array`
Alias: Array | | Empty Array | `array`
Alias: Array | `array`
Alias: Array | `array`
Alias: Array | | Null | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | Unidentified | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | Embedded MongoDB documents are inferred as Snowflake `VARIANT` values when read through the Spark MongoDB connector. Although Snowflake supports the `OBJECT` type, inferred embedded documents map to `VARIANT` unless explicitly declared otherwise. Learn more in [Supported data types](/data-analysis/gems/data-types). ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once a MongoDB connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # MSSQL Source: https://docs.prophecy.ai/data-analysis/environment/connections/mssql Learn how to connect with Microsoft SQL Server This page describes how to use and configure a connection to Microsoft SQL Server (MSSQL) in Prophecy. MSSQL is a relational database used for storing and querying structured data. ## Prerequisites Prophecy connects to Microsoft SQL Server (MSSQL) using the database credentials you provide. These credentials are used to authenticate your session and authorize all data operations during pipeline execution. To use an MSSQL connection effectively, your user account must have: * `select`, `insert`, `update`, and `delete` on the tables used in your Prophecy pipelines. * Access to the database and schema where tables are located. ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/external-table/mssql) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/external-table/mssql) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Index tables in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | ## Connection parameters To create a connection with Microsoft SQL Server, enter the following parameters: | Parameter | Description | | ------------------------------------------------------------------------ | --------------------------------------- | | Connection Name | Name to identify your connection | | Server | Address of the server to connect to | | Port | Port to use for the connection | | Username | Username for your MSSQL Server instance | | Password ([Secret required](/data-analysis/environment/secrets/secrets)) | Password for your MSSQL Server instance | ## Data type mapping When Prophecy processes data from Microsoft SQL Server (MSSQL) using SQL warehouses, it converts MSSQL-specific data types to formats compatible with your target warehouse. This table shows how [MSSQL data types](https://learn.microsoft.com/en-us/sql/t-sql/data-types/data-types-transact-sql?view=sql-server-ver17) are transformed for Databricks, BigQuery, and Snowflake. | MSSQL | Databricks | BigQuery | Snowflake | | ---------------- | --------------------------------- | --------------------------------- | --------------------------------- | | tinyint | `int`
Alias: Integer | `int64`
Alias: Integer | `number`
Alias: Number | | smallint | `int`
Alias: Integer | `int64`
Alias: Integer | `number`
Alias: Number | | int | `int`
Alias: Integer | `int64`
Alias: Integer | `number`
Alias: Number | | bigint | `bigint`
Alias: Bigint | `int64`
Alias: Integer | `number`
Alias: Number | | float / real | `double`
Alias: Double | `float64`
Alias: Float | `float`
Alias: Float | | decimal | `double`
Alias: Double | `float64`
Alias: Float | `number`
Alias: Number | | numeric | `double`
Alias: Double | `float64`
Alias: Float | `number`
Alias: Number | | money | `double`
Alias: Double | `float64`
Alias: Float | `number`
Alias: Number | | smallmoney | `double`
Alias: Double | `float64`
Alias: Float | `number`
Alias: Number | | bit | `boolean`
Alias: Boolean | `bool`
Alias: Boolean | `boolean`
Alias: Boolean | | char | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | varchar | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | text | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | nchar | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | nvarchar | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | ntext | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | xml | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | date | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | time | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | datetime | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | datetime2 | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | smalldatetime | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | datetimeoffset | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | rowversion | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | binary | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `binary`
Alias: Binary | | varbinary | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `binary`
Alias: Binary | | image | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `binary`
Alias: Binary | | uniqueidentifier | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `string`
Alias: String | | sql\_variant | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `string`
Alias: String | | geometry | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `binary`
Alias: Binary | | geography | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `binary`
Alias: Binary | | hierarchyid | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `binary`
Alias: Binary | Learn more in [Supported data types](/data-analysis/gems/data-types). ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once an MSSQL connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # Microsoft OneDrive Source: https://docs.prophecy.ai/data-analysis/environment/connections/onedrive Learn how to connect to OneDrive Microsoft OneDrive is a cloud-based file storage service that allows teams to store, access, and share files. In Prophecy, you can connect to OneDrive to read and write data as part of your data pipelines. ## Prerequisites To connect Prophecy to OneDrive, your Microsoft administrator must first [register Prophecy as an application](https://learn.microsoft.com/en-us/graph/auth/auth-concepts#register-the-application) in Microsoft Entra ID. This registration provides the Client ID and Client Secret needed to authenticate Prophecy with Microsoft APIs. As part of the setup, the following application-level permission must be granted to the registered app: * `Files.ReadWrite.All` This lets Prophecy read, create, update, and delete files in all site collections. Learn more in [Permissions for OneDrive API](https://learn.microsoft.com/en-us/onedrive/developer/rest-api/concepts/permissions_reference?view=odsp-graph-online). ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/file/onedrive) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/file/onedrive) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Index files in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | ## Connection parameters To create a connection with OneDrive, enter the following parameters. You can find the Tenant ID, Client ID, and Client Secret in your Microsoft Entra app. | Parameter | Description | | ----------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- | | Connection Name | Name to identify your connection | | Tenant ID | Your Microsoft Entra [tenant ID](https://learn.microsoft.com/en-us/entra/fundamentals/how-to-find-tenant) | | Client ID | Your Microsoft Entra app Client ID | | Client Secret ([Secret required](/data-analysis/environment/secrets/secrets)) | Your Microsoft Entra app Client Secret | | User Principal Name | Email you use to sign into Microsoft | ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once a OneDrive connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # Oracle DB Source: https://docs.prophecy.ai/data-analysis/environment/connections/oracle Learn how to connect to Oracle Oracle DB is a relational database management system. In Prophecy, you can connect to Oracle to read from and write to database tables as part of your pipelines. This page explains how to set up the connection, including required parameters, permissions, and how connections are shared within teams. ## Prerequisites Prophecy connects to Oracle using the database credentials you provide. These credentials are used to authenticate your session and authorize all data operations performed during pipeline execution. To use an Oracle connection effectively, your user account must have: * Read access to query data from tables * Write access to insert, update, or delete data ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/external-table/oracle) | Yes | | Write data with a Target gem | No | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Index tables in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | ## Connection parameters To create a connection with Oracle, enter the following parameters: | Parameter | Description | | ------------------------------------------------------------------------ | ---------------------------------------------------- | | Connection name | A name to identify your connection in Prophecy | | Server | Hostname of the Oracle database server | | Port | Port used by the Oracle database (default is `1521`) | | Username | Username for connecting to the Oracle database | | Database | Oracle Service Name or SID of the target database | | Password ([Secret required](/data-analysis/environment/secrets/secrets)) | Password for the specified user | ## Data type mapping When Prophecy processes data from Oracle using SQL warehouses, it converts Oracle-specific data types to formats compatible with your target warehouse. This table shows how [Oracle data types](https://docs.oracle.com/en/database/oracle/oracle-database/23/sqlrf/Data-Types.html) are transformed for Databricks, BigQuery, and Snowflake. | Oracle | Databricks | BigQuery | Snowflake | | ------------------------------- | ----------------------------------- | ----------------------------------------- | ---------------------------------- | | NUMBER | `decimal(38,5)`
Alias: Decimal | `bignumeric(38,5)`
Alias: BigNumeric | `number(38,10)`
Alias: Number | | SMALLINT / INTEGER / NUMBER(38) | `bigint`
Alias: Bigint | `int64`
Alias: Integer | `number`
Alias: Number | | FLOAT | `double`
Alias: Double | `float64`
Alias: Float | `float`
Alias: Float | | REAL / FLOAT(63) | `double`
Alias: Double | `float64`
Alias: Float | `float`
Alias: Float | | DOUBLE PRECISION / FLOAT(126) | `double`
Alias: Double | `float64`
Alias: Float | `float`
Alias: Float | | BINARY\_FLOAT | `double`
Alias: Double | `float64`
Alias: Float | `float`
Alias: Float | | BINARY\_DOUBLE | `double`
Alias: Double | `float64`
Alias: Float | `float`
Alias: Float | | DECIMAL | `decimal(38,5)`
Alias: Decimal | `bignumeric(38,5)`
Alias: BigNumeric | `number(38,10)`
Alias: Number | | NUMERIC | `decimal(38,5)`
Alias: Decimal | `bignumeric(38,5)`
Alias: BigNumeric | `number(38,10)`
Alias: Number | | CHAR | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | VARCHAR | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | VARCHAR2 | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | NCHAR | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | NVARCHAR2 | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | LONG | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | CLOB | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | NCLOB | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | BLOB | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `binary`
Alias: Binary | | DATE | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | TIMESTAMP | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | TIMESTAMP WITH TIME ZONE | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | TIMESTAMP WITH LOCAL TIME ZONE | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | `timestamp`
Alias: Timestamp | | INTERVAL YEAR TO MONTH | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | INTERVAL DAY TO SECOND | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | RAW | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `binary`
Alias: Binary | | LONG RAW | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `binary`
Alias: Binary | | XMLType | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | BOOLEAN | `boolean`
Alias: Boolean | `bool`
Alias: Boolean | `boolean`
Alias: Boolean | | URIType | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | DBURIType | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | XDBURIType | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | HTTPURIType | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | LongVarChar | `string`
Alias: String | `string`
Alias: String | `string`
Alias: String | | LongRaw | `binary`
Alias: Binary | `bytes`
Alias: Bytes | `binary`
Alias: Binary | For Oracle `number` types with explicit precision and scale, Prophecy preserves those values when mapping to Snowflake. For example, Oracle `number(10,2)` maps to Snowflake `number(10,2)`. Learn more in [Supported data types](/data-analysis/gems/data-types). ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once an Oracle connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # Postgres Source: https://docs.prophecy.ai/data-analysis/environment/connections/postgres Learn how to connect with PostgreSQL This page describes how to use and configure a connection to PostgreSQL in Prophecy. PostgreSQL is a relational database used for storing and querying structured data. ## Prerequisites Prophecy connects to PostgreSQL using the database credentials you provide. These credentials are used to authenticate your session and authorize all data operations during pipeline execution. To use a Postgres connection effectively, your user account must have: * `SELECT`, `INSERT`, `UPDATE`, and `DELETE` on the tables used in your Prophecy pipelines. * Access to the database and schema where tables are located. ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/external-table/postgres) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/external-table/postgres) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Index tables in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | Optional | ## Connection parameters To create a connection with PostgreSQL, enter the following parameters: | Parameter | Description | | ------------------------------------------------------------------------ | ------------------------------------------ | | Connection Name | Name to identify your connection | | Server | Address of the PostgreSQL server | | Port | Port to use for the connection | | Database | PostgreSQL database name | | Username | Username for your PostgreSQL instance | | Password ([Secret required](/data-analysis/environment/secrets/secrets)) | Password for your PostgreSQL instance | | Knowledge Graph Indexer | Enable or disable Knowledge Graph indexing | ## Data type mapping When Prophecy processes data from PostgreSQL using SQL warehouses such as Databricks or BigQuery, it converts PostgreSQL-specific data types to formats compatible with your target warehouse. This table shows common mappings and may vary depending on your configured warehouse. | PostgreSQL | Databricks | BigQuery | | ----------------- | ------------------------------- | ------------------------------- | | smallint | INT
Alias: Integer | INT64
Alias: Integer | | integer | INT
Alias: Integer | INT64
Alias: Integer | | bigint | BIGINT
Alias: Bigint | INT64
Alias: Integer | | decimal / numeric | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | real | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | double precision | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | boolean | BOOLEAN
Alias: Boolean | BOOL
Alias: Boolean | | char | STRING
Alias: String | STRING
Alias: String | | varchar | STRING
Alias: String | STRING
Alias: String | | text | STRING
Alias: String | STRING
Alias: String | | date | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | | time | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | | timestamp | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | | timestamptz | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | | bytea | BINARY
Alias: Binary | BYTES
Alias: Bytes | | uuid | STRING
Alias: String | STRING
Alias: String | | json / jsonb | STRING
Alias: String | STRING
Alias: String | Learn more in [Supported data types](/data-analysis/gems/data-types). ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once a Postgres connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # Microsoft Power BI Source: https://docs.prophecy.ai/data-analysis/environment/connections/power-bi Learn how to connect with PowerBI Prophecy supports writing data to Microsoft Power BI using the Power BI REST API. By configuring a connection with the appropriate Microsoft Entra credentials and scopes, you can push data directly from your Prophecy pipelines into tables used in reports and dashboards. ## Prerequisites To connect Prophecy to Power BI, your Microsoft administrator must first [register Prophecy as an application](https://learn.microsoft.com/en-us/graph/auth/auth-concepts#register-the-application) in Microsoft Entra ID. This registration provides the Client ID and Client Secret needed to authenticate Prophecy with Microsoft APIs. As part of the setup, the following scope must be granted to the registered app: * `Dataset.ReadWrite.All` This lets Prophecy update tables in Power BI. For detailed instructions on adding scopes to your app, visit [Using the Power BI REST APIs](https://learn.microsoft.com/en-us/rest/api/power-bi/#scopes) in the Power BI documentation. ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ---------------------------------------------------------------------------------------------- | --------- | | Read and write using [Source and Target gems](/data-analysis/gems/source-target/source-target) | No | | Write data with a [PowerBIWrite gem](/data-analysis/gems/report/power-bi) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | No | ## Limitations Prophecy uses the Power BI connection for the [Push Datasets](https://learn.microsoft.com/en-us/rest/api/power-bi/push-datasets) Power BI API. For the full list of limitations for this API, visit [Push semantic model limitations](https://learn.microsoft.com/en-us/power-bi/developer/embedded/push-datasets-limitations) in the Power BI documentation. ## Connection parameters To create a connection with Power BI, enter the following parameters. You can find the Tenant ID, Client ID, and Client Secret in your Microsoft Entra app. | Parameter | Description | | ----------------------------------------------------------------------------- | -------------------------------------- | | Connection Name | Unique name for the connection | | Tenant ID | Your Microsoft Entra tenant ID | | Client ID | Your Microsoft Entra app Client ID | | Client Secret ([Secret required](/data-analysis/environment/secrets/secrets)) | Your Microsoft Entra app Client Secret | ## Data type mapping Prophecy processes data using a SQL warehouse like Databricks SQL or BigQuery. When you are ready to write your transformed data to Power BI, data types are converted to [Power BI data types](https://learn.microsoft.com/en-us/power-bi/connect-data/desktop-data-types) using the following mapping. | Databricks | BigQuery | Power BI | | -------------------------------- | ------------------------------------------- | -------------- | | STRING
Alias: String | STRING
Alias: String | Text | | BOOLEAN
Alias: Boolean | BOOL
Alias: Boolean | True/False | | BYTE
Alias: Byte | INT64
Alias: Integer | Whole number | | SHORT
Alias: Short | INT64
Alias: Integer | Whole number | | INT
Alias: Integer | INT64
Alias: Integer | Whole number | | LONG
Alias: Long | INT64
Alias: Integer | Whole number | | FLOAT
Alias: Float | FLOAT64
Alias: Float | Decimal number | | DOUBLE
Alias: Double | FLOAT64
Alias: Float | Decimal number | | DECIMAL(p,s)
Alias: Decimal | NUMERIC/DECIMAL
Alias: Numeric/Decimal | Decimal number | | DATE
Alias: Date | DATE
Alias: Date | Date/Time | | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | Date/Time | | BINARY
Alias: Binary | BYTES
Alias: Bytes | Text (Base64) | | ARRAY
Alias: Array | REPEATED
Alias: Repeated | Text | | MAP\
Alias: Map | RECORD
Alias: Record | Text | | STRUCT
Alias: Struct | RECORD
Alias: Record | Text | | NULLTYPE
Alias: Nulltype | NULL
Alias: Null | Blank | ## Sharing connections within teams Power BI connections are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams in Prophecy. Once a Power BI connection is added to a fabric, anyone on that team can use it to send data to Power BI from their pipelines. Everyone will inherit the permissions of the Microsoft Entra app used for connection setup. # ProphecyManaged Source: https://docs.prophecy.ai/data-analysis/environment/connections/prophecy-managed Learn about the default ProphecyManaged fabric Available for [Free and Professional Editions](/data-analysis/administration/platform/editions) only. When you first sign in to Prophecy (Free or Professional Edition), you get access to an automatically-provisioned [fabric](/data-analysis/environment/fabrics/prophecy-fabrics). Each fabric includes a default SQL warehouse connection called **ProphecyManaged**. This connects you to [Prophecy In Memory](/data-analysis/environment/fabrics/prophecy-fabrics#supported-primary-sql-warehouses). This connection provides a ready-to-use environment for running SQL pipelines without additional configuration. ## Key features Review the following features to understand how the ProphecyManaged connection works. | Feature | Description | | ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Compute & Database** | Each connection has its own allocated compute resources and database in the warehouse. Prophecy uses a DuckDB warehouse under the hood. | | **Default Schema** | Outputs are written to a schema named `default` by default. You can change this in the connection settings. Agents write to the default schema when creating target tables. | | **Configuration** | Connection details (host, port, credentials) are fixed and cannot be modified. | | **Knowledge Graph** | You can reindex the connection to update the metadata so the Agent has current table information. | Data cannot be transferred between different ProphecyManaged connections. To share data with other users, add them to the team that owns the fabric with this connection. # Amazon Redshift Source: https://docs.prophecy.ai/data-analysis/environment/connections/redshift Learn how to connect to Redshift Connect Prophecy to your Amazon Redshift data warehouse to read from and write to tables from your pipelines. This page explains how to configure the connection, including required parameters, necessary permissions, and how connections are shared across teams. ## Prerequisites Prophecy connects to Amazon Redshift using the database credentials you provide. These credentials are used to authenticate your session and authorize all data operations performed during pipeline execution. To use a Redshift connection effectively, your user must have the following permissions: * `SELECT`, `INSERT`, `UPDATE`, and `DELETE` on the tables used in your Prophecy pipelines. * `CREATE TABLE`, `DROP TABLE`, or `ALTER TABLE` if your pipelines create or replace tables. * Access to specific schemas or databases where your tables reside. To learn more about user permissions, visit [Default database user permissions](https://docs.aws.amazon.com/redshift/latest/dg/r_Privileges.html) in the Amazon Redshift documentation. ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/external-table/redshift) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/external-table/redshift) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Index tables in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | ## Data type mapping When Prophecy processes data from Amazon Redshift using SQL warehouses, it converts Redshift-specific data types to formats compatible with your target warehouse. This table shows how [Amazon Redshift data types](https://docs.aws.amazon.com/redshift/latest/dg/c_Supported_data_types.html) are transformed for Databricks and BigQuery. | Redshift | Databricks | BigQuery | | ---------------- | --------------------------------- | ------------------------------- | | SMALLINT | INT
Alias: Integer | INT64
Alias: Integer | | INTEGER | BIGINT
Alias: Bigint | INT64
Alias: Integer | | BIGINT | BIGINT
Alias: Bigint | INT64
Alias: Integer | | REAL | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | DOUBLE PRECISION | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | DECIMAL | DECIMAL(38,5)
Alias: Decimal | NUMERIC
Alias: Numeric | | BOOLEAN | BOOLEAN
Alias: Boolean | BOOL
Alias: Boolean | | CHAR | STRING
Alias: String | STRING
Alias: String | | VARCHAR | STRING
Alias: String | STRING
Alias: String | | DATE | DATE
Alias: Date | DATE
Alias: Date | | TIME | TIMESTAMP
Alias: Timestamp | TIME
Alias: Time | | TIMETZ | TIMESTAMP
Alias: Timestamp | TIME
Alias: Time | | TIMESTAMP | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | | TIMESTAMPTZ | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | | VARBYTE | BINARY
Alias: Binary | BYTES
Alias: Bytes | | GEOMETRY | STRING
Alias: String | STRING
Alias: String | | GEOGRAPHY | STRING
Alias: String | STRING
Alias: String | | SUPER | STRING
Alias: String | STRING
Alias: String | | HLLSKETCH | STRING
Alias: String | STRING
Alias: String | | INTERVAL | STRING
Alias: String | STRING
Alias: String | Learn more in [Supported data types](/data-analysis/gems/data-types). ## Connection parameters To create a connection with Redshift, enter the following parameters: | Parameter | Description | | ------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------- | | Connection Name | Name to identify your connection | | Server | Redshift cluster server
Example: `redshift-cluster-1.abc123xyz789.us-west-2.redshift.amazonaws.com` | | Port | Port used by Redshift (default is `5439`) | | Username | Your Redshift username | | Database | Name of the Redshift database you want to connect to
Example: `analytics_db` | | Password ([Secret required](/data-analysis/environment/secrets/secrets)) | Your Redshift password | ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once a Redshift connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # Amazon S3 Source: https://docs.prophecy.ai/data-analysis/environment/connections/s3 Learn how to connect to Amazon S3 buckets Prophecy supports direct integration with Amazon S3, allowing you to read from and write to S3 buckets as part of your data pipelines. This page explains how to configure the connection, what permissions are required, and how S3 connections are managed and shared within your team. ## Prerequisites Prophecy connects to Amazon S3 using the AWS credentials you provide. These credentials are used to authenticate requests and authorize all file operations during pipeline execution. To ensure Prophecy can read from and write to S3 as needed, the credentials must grant the following permissions: * `s3:ListBucket` to list the contents of the bucket. * `s3:PutObject` to write files to the bucket. * `s3:GetObject` to read files. To learn more, visit [Required permissions for Amazon S3 API operations](https://docs.aws.amazon.com/AmazonS3/latest/userguide/using-with-s3-policy-actions.html) in the AWS documentation. ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | -------------------------------------------------------------------------------------------------------------------------------------- | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/file/s3) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/file/s3) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Trigger scheduled pipeline upon [file arrival or change](/data-analysis/production/scheduling/triggers#file-arrival-or-change-trigger) | Yes | | Index files in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | ## Connection parameters To create a connection with your Amazon S3 buckets, enter the following parameters: | Parameter | Description | | --------------------------------------------------------------------------------- | -------------------------------------------------------------------- | | Connection Name | Unique name for the connection | | Access Key ID | Your AWS access key ID | | Secret Access Key ([Secret required](/data-analysis/environment/secrets/secrets)) | Your AWS secret access key | | Region | AWS region where your S3 bucket is located
Example:` us-east-1` | | Bucket Name | Name of your S3 bucket | ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once an S3 connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # Salesforce Source: https://docs.prophecy.ai/data-analysis/environment/connections/salesforce Learn how to connect with Salesforce The Salesforce connection lets you access your [datasets](https://help.salesforce.com/s/articleView?id=analytics.bi_integrate_datasets.htm\&type=5) and [objects](https://developer.salesforce.com/docs/atlas.en-us.object_reference.meta/object_reference/sforce_api_objects_concepts.htm) from Salesforce in Prophecy. Prophecy leverages the Salesforce API to access and update your data. ## Prerequisites Prophecy connects to Salesforce using an API access token associated with your Salesforce account. Access to datasets and objects is controlled by the permissions granted to the account. Before setting up the connection, ensure your account has the necessary access to all relevant resources. For more details, visit [Dataset Security](https://help.salesforce.com/s/articleView?id=analytics.bi_integrate_dataset_security.htm\&type=5) and [Object Permissions](https://help.salesforce.com/s/articleView?id=platform.users_profiles_object_perms.htm\&type=5) in the Salesforce documentation. ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------- | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/external-table/salesforce) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/external-table/salesforce) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Index tables in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | ## Limitations You cannot browse your Salesforce datasets and objects in the Environment browser. Therefore, you cannot drag and drop tables from the Salesforce connection onto your canvas. Instead, all Salesforce Source and Target gems must be manually configured. ## Connection parameters To create a connection with Salesforce, enter the following parameters: | Parameter | Description | | ---------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | Connection Name | Unique name for the connection | | Salesforce URL | The base URL for your Salesforce instance.
Example: `https://yourcompany.my.salesforce.com` | | Username ([Secret required](/data-analysis/environment/secrets/secrets)) | Your Salesforce username used for authentication | | Password ([Secret required](/data-analysis/environment/secrets/secrets)) | Your Salesforce password used for authentication | | Access Token ([Secret required](/data-analysis/environment/secrets/secrets)) | The [Salesforce API access token](https://help.salesforce.com/s/articleView?id=xcloud.remoteaccess_access_tokens.htm\&type=5) associated with your account | | Salesforce API Version | The version of the Salesforce API to use | ## Data type mapping When Prophecy processes data from Salesforce using SQL warehouses, it converts Salesforce-specific data types to formats compatible with your target warehouse. This table shows how Salesforce data types are transformed for Databricks and BigQuery. The data types listed in the first column are the underlying API [Primitives](https://developer.salesforce.com/docs/atlas.en-us.object_reference.meta/object_reference/primitive_data_types.htm) and [Field Types](https://developer.salesforce.com/docs/atlas.en-us.object_reference.meta/object_reference/field_types.htm) (such as double, reference, and textarea) used by Salesforce developers and integration tools. This mapping reflects the technical storage format of the data. They are not the user-facing field names you see in the Salesforce setup menu (such as "Number," "Lookup Relationship," or "Long Text Area"). You can find mappings between these technical types and the corresponding user-interface names in the official Salesforce documentation. ### Primitive data types | Salesforce data type | Databricks | BigQuery | | -------------------- | -------------------------------- | ---------------------------------------- | | string | STRING
Alias: String | STRING
Alias: String | | boolean | BOOLEAN
Alias: Boolean | BOOL
Alias: Boolean | | double | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | double(p,s) | DECIMAL(p,s)
Alias: Decimal | NUMERIC / BIGNUMERIC
Alias: Numeric | | date | DATE
Alias: Date | DATE
Alias: Date | | dateTime | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | | time | STRING
Alias: String | TIME
Alias: Time | | base64 | BINARY
Alias: Binary | BYTES
Alias: Bytes | ### Data types for fields | Salesforce data type | Databricks | BigQuery | | -------------------- | ------------------------------------------------------ | ------------------------------------------------------ | | ID | STRING
Alias: String | STRING
Alias: String | | email | STRING
Alias: String | STRING
Alias: String | | percent | DECIMAL
Alias: Decimal | FLOAT64
Alias: Float | | phone | STRING
Alias: String | STRING
Alias: String | | currency | DECIMAL(16,2)
Alias: Decimal | FLOAT64
Alias: Float | | url | STRING
Alias: String | STRING
Alias: String | | encryptedstring | STRING
Alias: String | STRING
Alias: String | | picklist | STRING
Alias: String | STRING
Alias: String | | multipicklist | ARRAY\
Alias: Array | ARRAY\
Alias: Array | | reference | STRING
Alias: String | STRING
Alias: String | | location | STRUCT
Alias: Struct | STRUCT
Alias: Struct | | address | STRUCT
Alias: Struct | STRUCT
Alias: Struct | | textarea | STRING
Alias: String | STRING
Alias: String | | calculated | Depends on the [Formula Data Type](#calculated-fields) | Depends on the [Formula Data Type](#calculated-fields) | #### Calculated fields Calculated fields have distinct data types based on the [formula data type](https://help.salesforce.com/s/articleView?id=platform.choosing_a_formula_data_type.htm\&type=5) you select in Salesforce. | Formula data type | Databricks | BigQuery | | ----------------- | ------------------------------- | ------------------------------- | | Text | STRING
Alias: String | STRING
Alias: String | | Number | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | Checkbox | BOOLEAN
Alias: Boolean | BOOL
Alias: Boolean | | Date | DATE
Alias: Date | DATE
Alias: Date | | Date/Time | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once a Salesforce connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # SFTP Source: https://docs.prophecy.ai/data-analysis/environment/connections/sftp Learn how to set up SFTP in Prophecy SFTP (Secure File Transfer Protocol) is a secure way to transfer files over the internet using an encrypted connection between a client and a server. It's commonly used to exchange data between systems, especially in enterprise environments. In Prophecy, you can use an SFTP connection to read from and write to remote file systems directly in your data pipelines. This is useful when your data is stored outside cloud storage or databases. ## Prerequisites When you use an SFTP connection in Prophecy, permissions depend on the underlying SSH server and filesystem permissions on the server. Ensure you have the correct access to the files you need before setting up and using this connection. ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | -------------------------------------------------------------------------------------------------------------------------------------- | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/file/sftp) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/file/sftp) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Trigger scheduled pipeline upon [file arrival or change](/data-analysis/production/scheduling/triggers#file-arrival-or-change-trigger) | Yes | | Index files in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | ## Limitations Keep in mind the following limitations when using an SFTP connection. * **Simultaneous writes can cause file corruption.** If multiple processesβ€”such as different Prophecy jobsβ€”try to write to the same file at the same time using the same SFTP connection details, it can result in race conditions or corrupted files. This happens because the connector doesn't perform any client-side locking to coordinate access. * **Network latency affects transfer performance.** The speed and reliability of SFTP transfers depend on the physical distance between the SFTP server and Prophecy's infrastructure. Servers that are geographically closer to your Prophecy environment will generally provide faster, more stable performance. Servers located farther away may introduce higher latency, leading to slower or less consistent data transfers. For best results, use SFTP servers in the same region as your Prophecy environment. ## Connection parameters To configure an SFTP connection in Prophecy, enter the following parameters: | Parameter | Description | | --------------------- | ------------------------------------------------------------- | | Connection Name | Unique name for the connection | | Host | Hostname or IP address of the SFTP server | | Port | Port number for SFTP (default is `22`) | | Username | Your SFTP username | | Authentication Method | Choice between **Password** or **Private Key** authentication | ## Authentication methods You can configure your SFTP connection with one of the following authentication methods: * **Password:** Use a [secret](/data-analysis/environment/secrets/secrets) to enter your SFTP password. * **Private Key:** Upload a file that contains your SFTP private key. The file must be in [PEM format](https://en.wikipedia.org/wiki/Privacy-Enhanced_Mail) (`.pem` file). * There must be a header and a footer. * The content between the headers is a valid base64-encoded private key. * There are no extra spaces or newline characters. ## Supported ciphers Prophecy supports multiple encryption ciphers for SFTP connections to ensure both security and compatibility across different server configurations. Our SFTP client prioritizes modern, secure ciphers while maintaining fallback support for legacy systems. * `aes256-gcm@openssh.com` * `aes128-gcm@openssh.com` * `chacha20-poly1305@openssh.com` * `aes256-ctr` * `aes192-ctr` * `aes128-ctr` For compatibility with older SFTP servers, Prophecy also supports CBC legacy ciphers as fallbacks. We strongly recommend upgrading your SFTP servers to support GCM or CTR modes for enhanced security. ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once an SFTP connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # Microsoft SharePoint Source: https://docs.prophecy.ai/data-analysis/environment/connections/sharepoint Learn how to connect with SharePoint Prophecy supports integration with Microsoft SharePoint, allowing you to read from and write to SharePoint document libraries as part of your data pipelines. This connection enables you to work directly with files stored in SharePoint for processing, transformation, and reporting. ## Prerequisites To connect Prophecy to SharePoint, your Microsoft administrator must first [register Prophecy as an application](https://learn.microsoft.com/en-us/graph/auth/auth-concepts#register-the-application) in Microsoft Entra ID. This registration provides the Client ID and Client Secret needed to authenticate Prophecy with Microsoft APIs. As part of the setup, the registered app must be granted the `Sites.Selected` application permission, then explicitly given access to each SharePoint site the connection will use. `Sites.Selected` scopes access to specific site collections only, rather than granting tenant-wide access to every site. `Sites.Selected` requires additional setup beyond adding the permission in Entra ID β€” you must also grant access to each site individually via Microsoft Graph API. See [Grant service principal access to SharePoint sites](/data-analysis/gems/source-target/file/sharepoint#grant-service-principal-access-to-sharepoint-sites) for more details. Learn more in [Permissions for OneDrive and SharePoint API](https://learn.microsoft.com/en-us/onedrive/developer/rest-api/concepts/permissions_reference?view=odsp-graph-online). ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/file/sharepoint) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/file/sharepoint) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Index files in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | ## Limitations Prophecy can only access files stored in the [document library](https://support.microsoft.com/en-us/office/what-is-a-document-library-3b5976dd-65cf-4c9e-bf5a-713c10ca2872) of your SharePoint site. Make sure any files you want to import are placed there. Prophecy can only access sites that have been explicitly granted to it via `Sites.Selected`. If a connection can't reach a site, confirm the site has been granted access β€” see [Grant service principal access to SharePoint sites](/data-analysis/gems/source-target/file/sharepoint#grant-service-principal-access-to-sharepoint-sites). ## Connection parameters To create a connection with SharePoint, enter the following parameters. You can find the Tenant ID, Client ID, and Client Secret in your Microsoft Entra app. | Parameter | Description | | ----------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- | | Connection Name | Unique name for the connection | | Tenant ID | Your Microsoft Entra [tenant ID](https://learn.microsoft.com/en-us/entra/fundamentals/how-to-find-tenant) | | Client ID | Your Microsoft Entra app Client ID | | Client Secret ([Secret required](/data-analysis/environment/secrets/secrets)) | Your Microsoft Entra app Client Secret | | Site URL | URL of the SharePoint site to connect
Example: `https://yourcompany.sharepoint.com/sites/mysite` | The site entered here must match a site that has been granted access via `Sites.Selected`. See [Grant service principal access to SharePoint sites](/data-analysis/gems/source-target/file/sharepoint#grant-service-principal-access-to-sharepoint-sites). ## Sharing connections within teams Connections in Prophecy are stored in [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once you add a SharePoint connection to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is required: team members inherit access from the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissions, including access to sensitive data if applicable. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # Smartsheet Source: https://docs.prophecy.ai/data-analysis/environment/connections/smartsheet Learn how to connect with Smartsheet Smartsheet manages tasks, projects, and workflows using a spreadsheet-like interface that can contain rows, columns, and cell data. Prophecy uses the [Smartsheet API](https://developers.smartsheet.com/api/smartsheet/introduction) to establish the connection. ## Prerequisites Prophecy connects to Smartsheet using an API access token associated with your Smartsheet account. All operationsβ€”such as reading from or writing to sheetsβ€”are performed using the permissions granted to your user. To use a Smartsheet connection effectively, you must have the appropriate sharing permissions on the sheets you want to access. For example: * Viewer permission allows you to read sheet data but not modify it. * Editor permission is required to update or write data using a Target gem. Before setting up the connection, ensure your account has the necessary access to all relevant sheets. For more details, visit [Sharing permission levels](https://help.smartsheet.com/articles/1155182-sharing-permission-levels). ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/file/smartsheet) | Yes | | Write data with a [Target gem](/data-analysis/gems/source-target/file/smartsheet) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Index files in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | ## Limitations Keep in mind the following limitations when using the Smartsheet connection. * Smartsheet defines capacity limits in the [Limitations](https://developers.smartsheet.com/api/smartsheet/guides/basics/limitations) documentation. * Smartsheet allows users to create sheets with the same name in the same file location. Prophecy handles this situation in the following ways: * **Writing to Smartsheet when there are duplicate files.** In this case, the pipeline run will fail in Prophecy. Prophecy will not choose which file to overwrite with the Target gem. * **Reading from Smartsheet when there are duplicate files.** In this case, Prophecy appends a `(n)` to the duplicate files for differentiation in Prophecy. Duplicate Smartsheet file in Prophecy file browser ## Connection parameters To create a connection with Smartsheet, enter the following parameters: | Parameter | Description | | ---------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Connection Name | Unique name for the connection | | Access Token ([Secret required](/data-analysis/environment/secrets/secrets)) | Your [Smartsheet API access token](https://developers.smartsheet.com/api/smartsheet/guides/basics/authentication#access-token-best-practices) | ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once a Smartsheet connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # SMTP Source: https://docs.prophecy.ai/data-analysis/environment/connections/smtp Learn how to configure SMTP SMTP (Simple Mail Transfer Protocol) connections are used to send emails over the internet by allowing communication between email clients and servers. When you create an SMTP connection in Prophecy, the user credentials you provide are used to establish the connection. This user will always be the sender of an email when you use the Email gem with this connection. ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/source-target) | No | | Write data with an [Email gem](/data-analysis/gems/report/email) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | No | ## Limitations Only basic authentication is supported. The SMTP server must support plain username and password authentication. Prophecy does not currently support OAuth or other advanced authentication methods for SMTP connections. ## Connection parameters To create an SMTP connection, enter the following parameters: | Parameter | Description | | ------------------------------------------------------------------------ | ------------------------------------------------------------ | | Connection Name | Unique name for the connection | | URL | SMTP server URL
Example: `smtp.gmail.com` | | Port | SMTP port.
Port options may vary between SMTP services. | | Username | Your SMTP username | | Password ([Secret required](/data-analysis/environment/secrets/secrets)) | Your SMTP password | ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once an SMTP connection is added to a fabric, anyone in the team can use that connection to send emails from a pipeline. # Snowflake Source: https://docs.prophecy.ai/data-analysis/environment/connections/snowflake Learn how to connect with Snowflake Learn how to set up and use a Snowflake connection in Prophecy. With a Snowflake connection, you can read from and write to tables in your Snowflake account using Source and Target gems, browse data in the Environment browser, and process Snowflake data within Prophecy pipelines. Snowflake can also be configured as a fßabric SQL warehouse, where it executes pipeline transformations. This page describes how to configure a Snowflake connection for accessing data. To learn about using Snowflake as a compute engine, see [Create a Snowflake fabric](/data-analysis/environment/fabrics/create-fabrics/snowflake). ## Prerequisites When you create a Snowflake connection in Prophecy, all data operationsβ€”such as reading or writingβ€”are executed using the Snowflake credentials you provide. Ensure that your Snowflake user has the following permissions: * `SELECT`, `INSERT`, `UPDATE`, and `DELETE` on the tables used in your Prophecy pipelines. * `OWNERSHIP` on the table, if Prophecy needs to alter or replace it. When writing data through a Snowflake connection, Prophecy uploads Parquet files to a stage before loading them into Snowflake tables. This requires: * `CREATE FILE FORMAT` in the target schema. * `USAGE` on any file formats used for reading/writing Parquet files. * Write access to your user stage ## Connection type ### Connection type Snowflake can be used in two ways within Prophecy. | Connection type | Description | | ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **SQL Warehouse connection** | Snowflake acts as the compute engine for a fabric and executes pipeline SQL transformations. To learn more about SQL Warehouse connections, visit [Prophecy fabrics](/data-analysis/environment/fabrics/prophecy-fabrics). | | **Ingress/Egress connection** | Snowflake is used only as a data source or target. Pipelines read from or write to Snowflake tables while transformations run in another warehouse. | ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | SQL Warehouse | Ingress/Egress | | ---------------------------------- | ------------- | -------------- | | Run SQL queries | Yes | No | | Read data with Source gem | Yes | Yes | | Write data with Target gem | Yes | Yes | | Browse data in Environment browser | Yes | Yes | | Index tables in Knowledge Graph | No | No | ## Connection parameters To create a connection with Snowflake, enter the following parameters: | Parameter | Description | | --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | Connection Name | Name to identify your connection | | Account | URL of your Snowflake account
Example: `https://-.snowflakecomputing.com` | | Database | Default database for reading and writing data | | Schema | Default schema for reading and writing data | | Warehouse | Name of the Snowflake virtual warehouse used to execute queries for this connection. | | Role | Snowflake [role](https://docs.snowflake.com/en/user-guide/security-access-control-overview) of the user to connect
Example: `ACCOUNTADMIN` | | Authentication method | Enter your Snowflake username and use a [secret](/data-analysis/environment/secrets/secrets) to enter your password. | ## Data type mapping When Snowflake data is processed in pipelines running on other SQL warehouses (such as Databricks or BigQuery), Prophecy converts Snowflake-specific data types to compatible types. This table shows how [Snowflake data types](https://docs.snowflake.com/en/sql-reference/intro-summary-data-types) are transformed for Databricks and BigQuery. | Snowflake | Databricks | BigQuery | | -------------- | ------------------------------- | ------------------------------- | | NUMBER | BIGINT
Alias: Bigint | INT64
Alias: Integer | | INTEGER | BIGINT
Alias: Bigint | INT64
Alias: Integer | | BIGINT | BIGINT
Alias: Bigint | INT64
Alias: Integer | | SMALLINT | BIGINT
Alias: Bigint | INT64
Alias: Integer | | TINYINT | BIGINT
Alias: Bigint | INT64
Alias: Integer | | FLOAT | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | DOUBLE | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | REAL | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | DECIMAL | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | NUMERIC | DOUBLE
Alias: Double | FLOAT64
Alias: Float | | BOOLEAN | BOOLEAN
Alias: Boolean | BOOL
Alias: Boolean | | VARCHAR | STRING
Alias: String | STRING
Alias: String | | CHAR | STRING
Alias: String | STRING
Alias: String | | STRING | STRING
Alias: String | STRING
Alias: String | | TEXT | STRING
Alias: String | STRING
Alias: String | | DATE | DATE
Alias: Date | DATE
Alias: Date | | TIME | STRING
Alias: String | TIME
Alias: Time | | DATETIME | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | | TIMESTAMP\_NTZ | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | | TIMESTAMP\_LTZ | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | | TIMESTAMP\_TZ | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | | BINARY | BINARY
Alias: Binary | BYTES
Alias: Bytes | | VARBINARY | BINARY
Alias: Binary | BYTES
Alias: Bytes | | VARIANT | STRING
Alias: String | STRING
Alias: String | | OBJECT | STRING
Alias: String | STRING
Alias: String | | ARRAY | STRING
Alias: String | STRING
Alias: String | | NULL | STRING
Alias: String | STRING
Alias: String | Learn more in [Supported data types](/data-analysis/gems/data-types). ## Limitations There are a few limitations on the data types you can read from Snowflake: * Prophecy reads `Object`, `Array`, and `Variant` types as `String` type. * Prophecy does not support writing `Binary` type columns. ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once a Snowflake connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # Azure Synapse dedicated SQL pools Source: https://docs.prophecy.ai/data-analysis/environment/connections/synapse Learn how to connect with Azure Synapse Prophecy's Azure Synapse connector supports connecting to [dedicated SQL pools](https://learn.microsoft.com/en-us/azure/synapse-analytics/sql-data-warehouse/sql-data-warehouse-overview-what-is) that run Microsoft SQL Server (MSSQL). Use this connection instead of the [MSSQL connection](/data-analysis/environment/connections/mssql) when you host MSSQL in Azure Synapse dedicated SQL pool. This page describes how to set up the connection. ## Prerequisites Prophecy connects to Azure Synapse using the database credentials you provide. These credentials are used to authenticate your session and authorize all data operations during pipeline execution. To use an Azure Synapse connection effectively, your user account must have: * `SELECT`, `INSERT`, `UPDATE`, and `DELETE` on the tables used in your Prophecy pipelines. * Access to the database and schema where tables are located. ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/external-table/synapse) | Yes | | Write data with a Target gem | No | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | Yes | | Index tables in the [Knowledge Graph](/data-analysis/ai/knowledge-graph/knowledge-graph) | No | ## Connection parameters To create a connection with Azure Synapse, enter the following parameters: | Parameter | Description | | ------------------------------------------------------------------------ | -------------------------------------------------------------- | | Connection Name | Name to identify your connection | | Server | Address of the server to connect to | | Port | Port to use for the connection | | Username | Username for Synapse authentication | | Database | Name of the specific SQL database within your Synapse SQL pool | | Password ([Secret required](/data-analysis/environment/secrets/secrets)) | Password for Synapse authentication | ## Sharing connections within teams Connections in Prophecy are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams. Once a Synapse connection is added to a fabric, all team members who have access to the fabric can use the connection in their projects. No additional authentication is requiredβ€”team members automatically inherit the access and permissions of the stored connection credentials. Be mindful of the access level granted by the stored credentials. Anyone on the team will have the same permissionsβ€”including access to sensitive data if allowed. To manage this securely, consider creating a dedicated fabric and team for high-sensitivity connections. This way, only approved users have access to those credentials. # Tableau Source: https://docs.prophecy.ai/data-analysis/environment/connections/tableau Learn how to connect with Tableau Prophecy uses the [Tableau REST API](https://help.tableau.com/current/api/rest_api/en-us/REST/rest_api.htm) to send data to Tableau as `.hyper` files (Tableau's in-memory format). This page describes how to set up and use a Tableau connection, so you can publish and update data sources in your Tableau projects directly from Prophecy pipelines. ## Prerequisites To connect Prophecy to Tableau, you need to provide credentials in the form of a personal access token. These credentials are used to authenticate all actions performed via the Tableau REST API. To use a Tableau connection effectively, ensure that the personal access token has the necessary Publish capability for the Tableau project where you will publish data sources. For more details on Tableau permissions, see the Tableau documentation on [Permission Capabilities](https://help.tableau.com/current/server/en-us/permissions_capabilities.htm). ## Feature support The table below outlines whether the connection supports certain Prophecy features. | Feature | Supported | | ------------------------------------------------------------------------------------------ | --------- | | Read data with a [Source gem](/data-analysis/gems/source-target/source-target) | No | | Write data with a [TableauWrite gem](/data-analysis/gems/report/tableau) | Yes | | Browse data in the [Environment browser](/data-analysis/development/studio/studio#sidebar) | No | ## Limitations Using Tableau Hyper files is a legacy approach that Prophecy supports mainly for backward compatibility. It helps users keep existing dashboards running smoothly while migrating from older systems. For a modern, cloud-native workflow, write pipeline outputs directly to a supported cloud data platform like Databricks, Snowflake, or BigQuery. Then connect Tableau to that platform to visualize the dataβ€”no need to set up a separate Tableau connection or perform extra export steps in Prophecy. ## Connection parameters To create a connection with Tableau, enter the following parameters: | Parameter | Description | | ----------------------------------------------------------------------------- | ---------------------------------------------------------------------- | | Connection Name | Unique name for the connection | | Tableau Server URL | URL of your Tableau Server
Example: `https://tableau.example.com` | | Tableau Token Name | Name of your Tableau personal access token | | Tableau Token ([Secret required](/data-analysis/environment/secrets/secrets)) | Your Tableau personal access token | | Tableau Site Name | Name of the Tableau site you're connecting to | ## Data type mapping Prophecy processes data using a SQL warehouse like Databricks SQL or BigQuery. When you are ready to write your transformed data to Tableau, data types are converted to [Tableau data types](https://help.tableau.com/current/pro/desktop/en-us/datafields_typesandroles_datatypes.htm) using the following mapping. | Databricks | BigQuery | Tableau | | ------------------------------------ | ------------------------------------------------- | ---------------- | | BOOLEAN
Alias: Boolean | BOOL
Alias: Boolean | Boolean | | TINYINT
Alias: Tinyint | INT64
Alias: Integer | Number (whole) | | SMALLINT
Alias: Smallint | INT64
Alias: Integer | Number (whole) | | INT
Alias: Integer | INT64
Alias: Integer | Number (whole) | | BIGINT
Alias: Bigint | INT64
Alias: Integer | Number (whole) | | FLOAT
Alias: Float | FLOAT64
Alias: Float | Number (decimal) | | DOUBLE
Alias: Double | FLOAT64
Alias: Float | Number (decimal) | | DECIMAL(p,s)
Alias: Decimal | NUMERIC/BIGNUMERIC
Alias: Numeric/BigNumeric | Number (decimal) | | STRING
Alias: String | STRING
Alias: String | String | | BINARY
Alias: Binary | BYTES
Alias: Bytes | String | | DATE
Alias: Date | DATE
Alias: Date | Date | | TIMESTAMP
Alias: Timestamp | TIMESTAMP
Alias: Timestamp | Date & Time | | TIMESTAMP\_NTZ
Alias: Timestamp | DATETIME
Alias: Datetime | Date & Time | | INTERVAL
Alias: Interval | INTERVAL
Alias: Interval | String | | ARRAY
Alias: Array | ARRAY
Alias: Array | String | | STRUCT
Alias: Struct | STRUCT
Alias: Struct | String | | VOID
Alias: Void | NULL
Alias: Null | Null | ## Sharing connections within teams Tableau connections are stored within [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics), which are assigned to specific teams in Prophecy. Once a Tableau connection is added to a fabric, anyone on that team can use it to send data to Tableau from their pipelines. Everyone will inherit the permissions of the user authenticated during connection setup. # Windows network share Source: https://docs.prophecy.ai/data-analysis/environment/connections/windows-network-drive Synchronize files from a Windows network share to SharePoint so they can be used in Prophecy Organizations often store source data on Windows network shares or mapped drives. Because these file shares are not directly accessible from cloud-native applications, Prophecy recommends synchronizing them to Microsoft SharePoint Online and then [connecting Prophecy to SharePoint](/data-analysis/environment/connections/sharepoint). ## How it works Instead of connecting Prophecy directly to a Windows network share, use SharePoint Online as an intermediary. This approach allows you to continue managing files on your network while making them available through a supported cloud connection. ## Synchronize your network share Prophecy does not synchronize files between a network share and SharePoint. Instead, you should use a Microsoft-supported migration or synchronization tool to keep your SharePoint document library up to date. Microsoft recommends using one-way incremental synchronization, where changes made on the network share are periodically copied to SharePoint. This approach minimizes synchronization time while avoiding conflicts that can occur with two-way synchronization. Common Microsoft options include: * **[SharePoint Migration Tool (SPMT)](https://learn.microsoft.com/en-us/sharepointmigration/introducing-the-sharepoint-migration-tool)** for standalone or scripted migrations. * **\[SharePoint's Migration Manager]\([https://learn.microsoft.com/en-us/sharepointmigration/mm-get-started](https://learn.microsoft.com/en-us/sharepointmigration/mm-get-started)** for centralized enterprise deployments. For implementation details, refer to the [Microsoft SharePoint migration documentation](https://learn.microsoft.com/en-us/sharepointmigration/introducing-the-sharepoint-migration-tool). ## Connect Prophecy to SharePoint After your files are synchronized to SharePoint Online, [create a SharePoint connection in Prophecy](/data-analysis/environment/connections/sharepoint) that includes: * Your SharePoint site URL. * Your Azure Active Directory tenant ID. * Your Azure application (client) ID. * Your Azure client secret stored as a Prophecy secret. For the required connection properties, see the [Microsoft SharePoint connection properties API](https://learn.microsoft.com/en-us/sharepoint/dev/spfx/connect-to-sharepoint). ## Best practices * Use one-way incremental synchronization instead of two-way synchronization. * Schedule synchronization based on how frequently your source data changes. * Treat SharePoint as the data source that Prophecy reads. * After migration is complete, consider transitioning users to work directly from SharePoint rather than the original network share. # Create a Google BigQuery fabric Source: https://docs.prophecy.ai/data-analysis/environment/fabrics/create-fabrics/bigquery Connect Prophecy to your Google BigQuery environment Available on the [Enterprise Edition](/data-engineering/administration/platform/editions) only. To create a fabric using Google BigQuery as the provider, follow the steps below. Begin by creating a new fabric entity in Prophecy. 1. From the left sidebar, click the **+** sign. 2. On the Create Entity page, select **Fabric**. The **Basic Info** tab lets you define the key identifiers of the fabric. 1. Provide a name for the fabric. 2. (Optional) Provide a description for the fabric. 3. Select a [team](/data-analysis/administration/management/teams/teams) that can access the fabric. 4. Click **Continue**. The **Providers** tab allows you to choose a SQL warehouse provider. You cannot change the fabric provider after you create the fabric. 1. Under **Provider**, select **BigQuery**. The **Connections** tab allows you to store your credentials to various external data providers for reuse while attached to the fabric. * Under **SQL Warehouse Connection**, click **+ Connect SQL Warehouse**. * Select **BigQuery** as the connection type and configure the [connection details](/data-analysis/environment/connections/bigquery). You can add [Data Ingress/Egress Connections](/data-analysis/environment/connections/connections) to your fabric if you want, but these are not required. Click **Complete** to save the fabric. ## Column name handling BigQuery does not allow certain special characters in column names. When a pipeline uses a BigQuery fabric, Prophecy automatically encodes column names that contain unsupported characters for storage in BigQuery, then decodes them back to their original form for display throughout the pipeline UI and when writing to external targets. For example, when Prophecy encounters a column named `revenue(USD)`, it stores this column as `_PCYBQ__revenue_lp_USD_rp_` for BigQuery fabrics. In schema panels, the expression builder, and data previews, Prophecy displays the original name. You can hover over an encoded column name to see the physical name stored in BigQuery. The following characters trigger encoding: `!`, `"`, `$`, `(`, `)`, `*`, `,`, `.`, `/`, `;`, `?`, `@`, `[`, `\`, `]`, `^`, `` ` ``, `{`, `}`, `~`. See [BigQuery's docs](https://docs.cloud.google.com/bigquery/docs/schemas#flexible-column-names) for more information on this topic. Automatic encoding applies to column names ingested from external sources. Manually entering unsupported characters in column name fields (for example, in a Reformat gem) is not currently supported. ## Edit an existing fabric To edit an existing fabric: 1. From the left sidebar, click the **Metadata** page. 2. Navigate to the **Fabrics** tab. 3. Open the fabric you want to edit. 4. Update the fabric. Prophecy automatically saves your changes. ### Edit advanced settings To edit advanced settings, click the Advanced tab for the fabric. Here, you can select or deselect **Allow agent to access data**. This allows Prophecy agents to suggest data transformations and actions based on the data in the fabric. # Create a Databricks SQL fabric Source: https://docs.prophecy.ai/data-analysis/environment/fabrics/create-fabrics/databricks Connect Prophecy to your Databricks SQL warehouse Available for [Express and Enterprise Editions](/data-engineering/administration/platform/editions) only. To create a fabric using Databricks as the provider, follow the steps below. Begin by creating a new fabric entity in Prophecy. 1. From the left sidebar, click the **+** sign. 2. On the Create Entity page, select **Fabric**. The **Basic Info** tab lets you define the key identifiers of the fabric. 1. Provide a name for the fabric. 2. (Optional) Provide a description for the fabric. 3. Select a [team](/data-analysis/administration/management/teams/teams) that can access the fabric. 4. Click **Continue**. The **Providers** tab allows you to configure the SQL warehouse provider you want to use. You cannot change the fabric provider after you create the fabric. 1. Under **Provider**, select **Databricks**. 2. Enter your Databricks Workspace URL. Example: `https://dbc-.cloud.databricks.com`. 3. Select the authentication method you want to use. For detailed instructions, see [Authentication methods](/data-analysis/environment/connections/databricks#authentication-methods). 4. Select the **SQL Warehouse** checkbox. 5. Enter the **HTTP Path** for your[ Databricks SQL warehouse](https://docs.databricks.com/aws/en/integrations/compute-details). Example: `/sql/1.0/warehouses/`. 6. Provide the Databricks **Catalog** and **Schema** where Prophecy will write data to by default. 7. Optionally, select the **Compute Cluster** checkbox and select the cluster you want to use for any Spark operations. Prophecy runs Spark on [Databricks serverless](/data-engineering/fabrics/spark-provider/databricks/databricks-serverless) when no compute cluster is selected. Click **Complete** to save the fabric. To add [secrets](/data-analysis/environment/secrets/secrets) and [connections](/data-analysis/environment/connections/connections) to your fabric, you need to save the fabric first. ## Edit an existing fabric To edit an existing fabric: 1. From the left sidebar, click the **Metadata** page. 2. Navigate to the **Fabrics** tab. 3. Open the fabric you want to edit. 4. Update the fabric. Prophecy automatically saves your changes. ### Edit advanced settings To edit advanced settings, click the Advanced tab for the fabric. Here, you can select or deselect **Allow agent to access data**. This allows Prophecy agents to suggest data transformations and actions based on the data in the fabric. # Create a Snowflake fabric Source: https://docs.prophecy.ai/data-analysis/environment/fabrics/create-fabrics/snowflake Learn how to connect to Snowflake A Snowflake connection allows Prophecy to access tables in your Snowflake account and execute queries using a Snowflake virtual warehouse. This page explains how to use and configure a Snowflake connection in Prophecy. Available for [Express and Enterprise Editions](/data-engineering/administration/platform/editions) only. ## Prerequisites Before creating the Snowflake fabric, ensure you have: * A Snowflake account URL. * A Snowflake warehouse to execute queries. * A role with permissions to use the warehouse and access the database, schema, and stage. * A stage with read and write permissions. ## Creating a Snowflake fabric To create a fabric using Snowflake as the provider, follow the steps below. Begin by creating a new fabric entity in Prophecy. 1. From the left sidebar, click the **+** sign. 2. On the Create Entity page, select **Fabric**. The **Basic Info** tab lets you define the key identifiers of the fabric. 1. Provide a name for the fabric. 2. (Optional) Provide a description for the fabric. 3. Select a [team](/data-analysis/administration/management/teams/teams) that can access the fabric. 4. Click **Continue**. The **Providers** tab allows you to configure the SQL warehouse provider you want to use. You cannot change the fabric provider after you create the fabric. 1. Under **Provider**, select **Snowflake**. 2. Enter your **Account URL**. Example: `https://xy12345.us-east-1.snowflakecomputing.com` 3. Enter the **Role** that Prophecy will use to access Snowflake resources. Example: `TRANSFORM_ROLE` 4. Enter the **Warehouse** that will execute SQL queries generated by your pipelines. Example: `ANALYTICS_WH` 5. Provide the default **Database** and **Schema** where Prophecy will read and write data. Example: `Database: ANALYTICS` `Schema: PIPELINES` 6. Enter the **Prophecy Dataplane URL**. (This is used to orchestrate pipeline execution and data movement.) 7. Select the authentication method you want to use. **Username and Password** (typical for user credentials) * Enter the **Username**. * Enter the **Password**. **Key-pair** (authenticate with an RSA key pair instead of a password) * Enter the **Username**. * Upload your **Private Key** file. This method does not currently support passphrase-protected private keys β€” there's no field to enter a passphrase. Use an unencrypted private key, or choose a different authentication method. **OAuth** (used when authentication is managed by an identity provider or service principal.) * Select an **App Registration** from the dropdown. This lists the OAuth app registrations your cluster admin has configured for Snowflake, including native Snowflake OAuth registrations and any [ID Anywhere (Snowflake)](/administration/management/cluster-admin-settings/oauth-setup) registrations, which authenticate through an ADFS/ID Anywhere identity provider instead of Snowflake's native OAuth. If the registration you need isn't listed, ask your cluster admin to add it under **Settings** > **Admin** > **Security**. 8. Provide the **Stage** that Prophecy will use for temporary orchestration files. Example: `@PROPHECY_STAGE` The stage functions similarly to Databricks volumes. Prophecy uses it to store temporary files required during pipeline orchestration. The configured role must have **read and write permissions** on the stage. Click **Complete** to save the fabric. To add [secrets](/data-analysis/environment/secrets/secrets) and [connections](/data-analysis/environment/connections/connections) to your fabric, you need to save the fabric first. ## Edit an existing fabric To edit an existing fabric: 1. From the left sidebar, click the **Metadata** page. 2. Navigate to the **Fabrics** tab. 3. Open the fabric you want to edit. 4. Update the fabric. Prophecy automatically saves your changes. ### Edit advanced settings To edit advanced settings, click the Advanced tab for the fabric. Here, you can select or deselect **Allow agent to access data**. This allows Prophecy agents to suggest data transformations and actions based on the data in the fabric. ## Limitations Currently, Snowflake fabrics do not support: * Case-sensitive identifiers. * Creating new partitioned tables. * Modifying partitioning of existing tables. # What is a fabric? Source: https://docs.prophecy.ai/data-analysis/environment/fabrics/prophecy-fabrics Fabrics let you connect to external compute and data A fabric is a Prophecy entity that contains all of the connection information you need to connect to **external compute** and **data storage**. You need fabrics to run pipelines. Prophecy provides you with compute and data storage for Free and Professional Editions. Users on Express and Enterprise Editions need to connect to their own compute and data storage. ## Key concepts ### External compute A fabric connects to a SQL engine that runs your pipeline transformations. This compute engine executes the SQL queries and performs the data transformations that your pipelines generate. You access compute through your primary SQL warehouse connection, such as Databricks SQL, BigQuery, Snowflake, or Prophecy In Memory. ### Data storage A fabric connects to storage systems that hold your data. You read from sources and write to targets in your pipelines. There are two types of data storage: * **Warehouse data storage**: Storage that is integrated with your compute engine. For example, when you use a Databricks SQL warehouse connection, it accesses data in Unity Catalog. SQL warehouses can only access data that is local to the warehouse. They cannot read from or write to external systems. * **External data storage**: Storage systems that are separate from your compute engine. To move data in and out of the warehouse, you use ingress/egress connections through [Prophecy Automate](/data-analysis/administration/platform/architecture#what-is-prophecy-automate) (Prophecy-native runtime). Prophecy Automate handles data movement between external systems and your warehouse. This includes managing secrets and credentials necessary to access external systems. ## Fabric components The following table describes how fabric components map to the key concepts. | Component | Description | | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | SQL Warehouse Connection | Connection to the SQL compute engine and the coupled data storage for that provider. A primary warehouse connection is required for every Prophecy fabric to execute SQL queries generated by your pipeline. | | Ingress/Egress Connections | Connections to external data providers. Each connection can be used to read and write data from supported locations. For the full list of supported connections, visit [Connections](/data-analysis/environment/connections/connections). | | Secrets | Secrets store sensitive data. Use Prophecy secrets to store credentials that authenticate your connections in the fabric. To learn more, see [Secrets](/data-analysis/environment/secrets/secrets). | ```mermaid theme={null} flowchart TD A[Fabric] --> C[SQL Warehouse Connection] A --> D[Connections] A --> E[Secrets] C --> F[Compute] C --> G[Data Storage] D --> H[Data Storage] ``` ## Fabric provisioning Prophecy fabrics are provisioned differently depending on your [Prophecy edition](/data-analysis/administration/platform/editions). | Edition | Description | | ---------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Free & Professional Editions | Each user automatically receives a personal fabric that is linked to their team. These fabrics run on **Prophecy In Memory**, and all execution is fully managed by Prophecy. Usage is billed in credits. | | Express Edition | Users create fabrics manually and connect them to their own Databricks SQL warehouse. Prophecy Automate is included, while Databricks costs are managed separately by the user. | | Enterprise Edition | Users create fabrics manually and can connect to any external SQL engine supported by Prophecy, including Databricks SQL warehouse, BigQuery, and Snowflake. Prophecy Automate is included, and warehouse costs are managed separately by the user. | Free and Professional Editions do not support creating new fabrics. ## Supported SQL warehouses While the Profession Edition of Prophecy includes built-in compute, Enterprise users need to connect Prophecy to their own compute environment. Prophecy supports multiple compute providers and adapts to each provider's capabilities: * Some providers support SQL execution only. * Some providers support Spark execution only. * Some providers support both SQL and Spark. For **data analysis** in SQL, we support the following SQL warehouse providers: | SQL Warehouse | Description | | ------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Prophecy In Memory | Exclusive to Free and Professional Editions. This warehouse is fully managed by Prophecy and powered by DuckDB. All pipeline transformations execute using the DuckDB SQL dialect. | | [Databricks SQL](/data-analysis/environment/fabrics/create-fabrics/databricks) | Supported in Express and Enterprise Editions. Connect your own Databricks SQL warehouse to run pipeline transformations. | | [Google BigQuery](/data-analysis/environment/fabrics/create-fabrics/bigquery) | Supported in Enterprise Edition. Connect to your BigQuery environment to run pipeline transformations. | | [Snowflake](/data-analysis/environment/fabrics/create-fabrics/snowflake) | Supported in Enterprise Edition. Connect to your Snowflake environment to run pipeline transformations. | The warehouse you choose provides the compute and determines the SQL dialect for all pipelines in the fabric. Each SQL warehouse has different capabilities. For example, Snowflake fabrics currently do not support partitioned table operations or case-sensitive identifiers. Visit each provider page to learn more about how to connect to the provider and configure the fabric. # Prophecy secrets Source: https://docs.prophecy.ai/data-analysis/environment/secrets/secrets Use the Prophecy-native secret manager A **secret** is stored sensitive data such as passwords, API keys, or certificates. Prophecy includes a built-in secret manager to store secrets securely and allow [Prophecy Automate](/data-analysis/administration/platform/architecture) to access them at runtime. This prevents credentials from being hardcoded or shared in plain text. | Secret type | Description | Example use case | | ------------------- | ----------------------------------------------------------------- | --------------------------------- | | Text | A string value. | Access token value | | Binary | A file that you upload. | SSL certificate | | Username & Password | Two-field credential for a username and password. | Basic authentication for RestAPIs | | M2M OAuth | Multi-field credential used for client credential authentication. | REST APIs using bearer tokens | ## Access control Access to secrets is related to fabric access. * Secrets are tied to the fabric where they're created. * Anyone with access to that fabric can reference its secrets in projects. * The secret value is never visible, even to the user who created it. ## Create a secret To add a new secret to a fabric: 1. From the left sidebar, open the **Metadata** page. 2. Open the fabric where you want to store the secret. 3. Go to the **Secrets** tab. 4. Click **+ Add Secret**. 5. In the dialog, choose the **Secret Type**. | Secret Type | Parameters | | ------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Text |
  • **Name**: Label to identify the secret
  • **Value**: String value
| | Binary |
  • **Name**: Label to identify the secret
  • **Value**: Upload a file (such as a certificate)
| | Username & Password |
  • **Name**: Label to identify the secret
  • **Username**: Username for authentication
  • **Password**: Password for authentication
| | M2M OAuth |
  • **Name**: Label to identify the secret
  • **Client ID**: OAuth client identifier
  • **Client Secret**: OAuth client secret key
  • **Auth URL**: OAuth authorization server endpoint
  • **Scope** (Optional): OAuth permissions scopes
| 6. Click **Create** to save it to the fabric. ## Reference a secret Secrets are often used to hide credentials in [connections](/data-analysis/environment/connections/connections). 1. Create a new connection. 2. In fields that require credentials, a secret picker appears automatically. 3. If the secret doesn't exist yet, click the **New Secret** button, which opens the fabric's **Add Secret** dialog. 4. Once saved, the secret will appear in the picker. 5. Select it, and Prophecy will validate the connection securely using that secret. # Gem auto-documentation Source: https://docs.prophecy.ai/data-analysis/gems/copilot/auto-documentation Generate documentation throughout your pipeline to improve transparency As pipelines scale, they can become harder to interpret and maintain. To improve clarity and collaboration, Prophecy allows you to add labels and descriptions to gems, helping others understand the purpose of each step in the pipeline. To streamline this process, Copilot can automatically generate or update labels and descriptions for you. ## Understand your data Before diving into individual pipeline steps, it helps to understand the data itself. Copilot can automatically generate: * **Table descriptions** for Source and Target gems * **Column descriptions** (metadata) to provide field-level context These descriptions offer a high-level view of your data, helping you and others quickly grasp what data the pipeline is working with. Generate source and target descriptions ## Understand each gem When you want to know a specific gem's purpose in the pipeline, you can ask Copilot to explain it in plain language. Find the **Explain** option in the gem action menu. To learn more, visit the documentation on [gems](/data-analysis/gems/gems). ## Improve gem labels Clear labels make pipelines easier to read. Copilot will automatically generate gem labels as you build you pipeline. If you update a gem, Copilot will ask for permission update the gem label. Generate gem label You also have the option to add comments to gems. While Copilot doesn't generate comments, you can still add them manually as tooltips to share helpful notes or warnings with others. Find the **Label** and **Add Comment** options in the gem action menu. To learn more, visit the documentation on [gems](/data-analysis/gems/gems). # Generate expressions in gems Source: https://docs.prophecy.ai/data-analysis/gems/copilot/gem-expressions Automatically generate expressions with natural language Copilot enhances the expression-building experience within gem configurations in two ways: * Suggesting context-aware expressions based on your input and column metadata. * Accepting natural language prompts to build new expressions. ## Copilot suggestions Copilot suggests expressions based on semantic cues from your column names. To accept Copilot suggestions, click `tab` or click on the expression itself in the gem. For best results, use descriptive column names. * **Recommended**: `customer_email` Clear intent, likely to produce accurate suggestions. * **Not recommended**: `col1` Ambiguous, may result in less reliable output. Copilot expressions ## Copilot prompts If you want to build your own expression, you can ask Copilot to help. Copilot is available for writing expressions in both the visual and code view. 1. Click on an expression field in a gem. 2. Click **Ask AI** to enter a prompt. 3. Copilot uses your prompt to generate an expression. ## Example prompts The following are example prompts you can use to generate expressions in gems. | Scenario | Prompt | | -------------------- | ------------------------------------------------------------------ | | Basic transformation | β€œExtract the domain from the `email` column.” | | String formatting | β€œCapitalize the first letter of each word in `customer_name`.” | | Date parsing | β€œConvert `order_date` to YYYY-MM format.” | | Conditional logic | β€œIf `amount` > 1000, then label as 'high value', else 'standard'.” | | Null handling | β€œReplace null values in `zipcode` with '00000'.” | | Math calculation | β€œCalculate discount as `price` \* `discount_rate`.” | # Error fixing Source: https://docs.prophecy.ai/data-analysis/gems/copilot/gem-fixes Fix gems errors with one click Troubleshooting pipeline errors manually can be time-consuming. Prophecy simplifies this process by offering multiple ways to detect and resolve issues quickly. After Copilot makes the fix, Prophecy will always provide a summary of the changes made. ## Inline fix Click an expression inside a gem with an error and select **Fix it** to apply an automatic correction. Copilot inline fix ## Fix on save If you try to save a gem with detected errors, a **Fix Diagnostics** dialog appears. * **Save without fixing**: Save the gem without fixing the errors. * **Fix diagnostics**: Let Copilot fix the errors automatically. Copilot fix on save ## Gem-level fix On the pipeline canvas, open a gem's action menu and select **Fix** to automatically correct errors within that component. Copilot fix gem # Condition Source: https://docs.prophecy.ai/data-analysis/gems/custom/condition Route an input dataset to one of multiple outputs based on ordered conditions. The Condition gem evaluates a series of conditions and routes the entire input dataset to the first output whose condition evaluates to `true`. This makes it useful for pipeline control flow, data quality checks, and conditional execution. ## How it works Each output port contains a condition that must evaluate to a single boolean value (`true` or `false`). The Condition gem evaluates output ports from top to bottom: 1. Evaluate the first output condition. 2. If the condition evaluates to `true`, route the entire input dataset to that output. 3. Otherwise, evaluate the next output. 4. Continue until a condition evaluates to `true`. 5. If no conditions evaluate to `true`, route the dataset to the final (default) output. Only one output receives the input dataset. The Condition gem chooses a single route for the entire dataset; it does not split rows across multiple outputs. After a condition evaluates to `true`, later conditions are not evaluated. ## Condition requirements Each condition must return exactly one row containing a boolean value. Conditions that return multiple rows are invalid and cause the pipeline to fail with a scalar subquery error. For example: | Condition | Result | Valid | | ------------------------- | ------------------------ | ----- | | `count > 1000` | One boolean value | βœ“ | | `sum(amount) > threshold` | One boolean value | βœ“ | | `A = 1` | One result per input row | βœ— | Conditions typically use aggregate functions or expressions that evaluate the input dataset as a whole. ## Use the Condition gem 1. Add a **Condition** gem to your pipeline from the **Custom** category. 2. Connect an input to the gem. 3. Define a condition for each output. 4. Click **Add Routing Rule** or **+** to add additional outputs. 5. Arrange outputs in the order you want them evaluated. 6. Connect downstream gems to each output. The gem starts with two output ports by default. ## Output behavior Only one output receives the input dataset. All downstream branches still execute, even when they receive zero rows. Outputs that are not selected receive empty dataframes. ```mermaid theme={null} flowchart LR A[Input dataset] --> B{out0 condition
returns TRUE?} B -->|Yes| C[Route entire dataset
to out0] B -->|No| D{out1 condition
returns TRUE?} D -->|Yes| E[Route entire dataset
to out1] D -->|No| F[Route entire dataset
to default output] ``` For example, if `out0` evaluates to `true`, the pipeline behaves like this: | Output | Rows | | ------ | -------------- | | out0 | All input rows | | out1 | 0 | | out2 | 0 | Downstream transformations should therefore handle empty dataframes. Some operations, such as aggregations, may still produce output when their input dataframe is empty. For example, a count aggregation can return a single row. ## Example: Route based on row count Assume your pipeline receives 25 rows. Configure the Condition gem as follows: * `out0`: `count(*) < 10` * `out1`: `count(*) < 100` * `out2`: default The Condition gem evaluates the outputs in order. * `count(*) < 10` evaluates to `false`. * `count(*) < 100` evaluates to `true`. The entire dataset is routed to `out1`. `out0` and `out2` still execute but receive empty dataframes. ## Example: Invalid condition Assume the input contains: | A | | - | | 1 | | 2 | The following condition is invalid: ``` A = 1 ``` The expression produces two results: | Result | | ------ | | true | | false | Because the condition returns multiple rows instead of a single boolean value, the Condition gem fails with a scalar subquery error. To use the Condition gem successfully, each condition must evaluate to exactly one boolean value. ## Example use cases Use the Condition gem when you need to: * Route a pipeline based on dataset size. * Apply data quality guardrails. * Trigger different processing paths based on aggregate metrics. * Implement pipeline control flow. ## Limitations * Conditions are evaluated from top to bottom. * The first condition that evaluates to `true` determines the output. * Only one output receives the input dataset. * Each condition must evaluate to exactly one boolean value. * Conditions that return multiple rows cause the pipeline to fail. * Downstream logic must handle empty dataframes. ## Filter vs Condition gem Use the [Filter gem](/data-analysis/gems/prepare/filter) when you want to keep or remove rows. Use the Condition gem when you want to choose between multiple execution paths. | Filter gem | Condition gem | | ----------------------------- | ---------------------------------------- | | Filters rows | Chooses one pipeline branch | | Output contains matching rows | Output contains the entire input dataset | | Evaluates each row | Evaluates one boolean condition | | Can reduce a dataset | Selects one output path | # Directory gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/custom/directory List files and folders of a specified directory This gem runs in . ## Overview List files and folders of a specified directory from a data ingress/egress connection. ## Prerequisites * Run Prophecy version 4.1.3 or higher. ## Input and Output The Directory gem does not accept any inputs. The Directory gem produces one output. The output schema includes the following columns: * `name`: The name of the file. * `path`: The full path to the file. * `size_in_bytes`: The size of the file. Folders will be listed as `0` bytes. * `creation_time`: The time that the file was created. * `modification_time`: The time that the file was last modified. * `parent_directory`: The parent directory of the file or folder. * `file_type`: Whether the record listed is a file or a folder. * `sheet_name`: The name of the excel sheet in an XLSX file. This column appears when the **Include sheet name as column in output for xlsx files** parameter is enabled. If a certain connection does not provide a certain field (for example, Databricks does not provide creation time), then the columns will be populated with zeroes or null values. ## Parameters Configure the Directory gem using the following parameters. | Parameter | Description | | ----------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Connection type | The data provider to connect to. See [Supported connection types](#supported-connection-types). | | Select or create connection | New or existing connection to the provider you selected. | | Path | Path to directory that you want to see the contents of. | | Enable to include files/directories inside subfolders | When enabled, the gem recursively traverses and include all files and directories within subdirectories of the specified path. | | File pattern (Optional) | Regular expression (regex) pattern used to narrow results to matching entries. | | Include sheet name as column in output for xlsx files | When enabled, the gem adds a column to the output that includes XLSX sheet names. If a file has multiple sheets, one row is generated per sheet name. For example, a file with three sheets produces three rowsβ€”one for each sheet. This field is `null` if the file is not an XLSX file. | ## Supported connection types You can use the Directory gem to list files and folders from the following connection types: * [Databricks Volumes](/data-analysis/environment/connections/databricks) * [Amazon S3](/data-analysis/environment/connections/s3) * [OneDrive](/data-analysis/environment/connections/onedrive) * [SFTP](/data-analysis/environment/connections/sftp) * [SharePoint](/data-analysis/environment/connections/sharepoint) * [Smartsheet](/data-analysis/environment/connections/smartsheet) # DynamicInput Source: https://docs.prophecy.ai/data-analysis/gems/custom/dynamic-input Run SQL queries that update dynamically at runtime This gem runs in . ## Overview Use the DynamicInput gem to run SQL queries or read data files, dynamically automating data retrieval. This page describes how to either [modify SQL queries](#read-option-1-modify-sql-query) or [dynamically read data from multiple files](#read-option-2-dynamically-read-data-from-multiple-files). ## Prerequisites * Run Prophecy version 4.1.3 or higher. ## Read Option 1: Modify SQL Query Run parameterized SQL queries that automatically update based on incoming data. Define a single SQL template with placeholders, and at runtime, DynamicInput replaces those placeholders with column values from your input dataset. ### Input and Output Configure the following input and output ports. | Port | Description | | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | **in0** | Input dataset that contains the placeholder replacement values to be used in the SQL query template.
Each row in this dataset generates its own SQL query. | | **out** | Output dataset containing the combined results from all generated queries.
The gem executes one query per input row, then unions the results into a single output. | ### Parameters Configure the DynamicInput gem using the following parameters. | Parameter | Description | | --------------------------- | ---------------------------------------------------------------------------------------------------------------- | | Select Read Options | The **Modify SQL Query** option denotes that the gem reads data using a dynamically modified SQL query. | | Table Connection Type | Select the type of [data ingress/egress connection](/data-analysis/environment/connections/connections) to use. | | Select or create connection | Choose an existing connection or create a new one for the selected source. | | Table or Query | Write a SQL query template that will be used to retrieve data. Can include placeholders for dynamic replacement. | | Pass fields to the Output | Select specific columns that contain the data you will use to replace strings. | | Replace a Specific String | Reference a static text value to replace in the query, and select the column with replacement values. | ### Example This example demonstrates how to filter rows from an Oracle table by replacing placeholder text with values from your dataset. 1. Under **Select Read Options**, select **Modify SQL Query**. 2. Under **Table Connection Type**, select **Oracle**. 3. Select an existing connection or [create a new one](/data-analysis/environment/connections/oracle). 4. In **Table or Query**, enter: ``` SELECT * FROM SALES.ORDERS WHERE status <> 'replace' AND region <> 'AAA' ``` 5. In the **Replace a Specific String** table, set: * **Text to Replace**: `replace` * **Replacement Field**: `status_column` When you run the gem, it will take each row from your input dataset, replace the placeholder text (`replace`) in the SQL template with the corresponding value from `status_column`, and execute that query. Prophecy will then union all of the query results into a single output dataset, so you can work with them as one combined table. ## Read Option 2: Dynamically read data from multiple files Combine data from listed XLSX files, or extract only the sheet names from listed files and append them as values in a new column. ### Input and Output Configure the following input and output ports. | Port | Description | | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | | **in0** | Input dataset that lists the file paths to the XLSX files to read. Each path in this dataset is processed according to the selected output mode. | | **out** | Returns either the unioned data from all XLSX files or a list of sheet names, depending on the selected output mode. | ### Parameters Configure the DynamicInput gem using the following parameters. | Parameter | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Select Read Options | Select **Dynamically read data from multiple files** to enable reading and combining XLSX files at runtime. | | Select output mode | Choose how to process the XLSX files.
  • **Union sheet data by column name**: Read and combine data from all XLSX files in the provided file paths. Empty or non-XLSX files are skipped.
  • **Retrieve sheet names**: Extract only the sheet names from each file and append them as values in a new column.
| | File Connection Type | Select the file storage [connection](/data-analysis/environment/connections/connections) type to use. | | Select or create connection | Choose an existing connection or create a new one for the selected connection type. | | File Path Column | Specify the column in the input dataset that contains the file paths to read from. | | Password | (Optional) Provide a password to access password-protected XLSX files. | The following additional parameters are applicable to the **Union sheet data by column name** option only. | Parameter | Description | | ----------------- | --------------------------------------------------------------------------------------------------------------------------------- | | Sheet Name Column | Specify the column in the input dataset that contains the sheet name to read from in each XLSX file. | | Header Row | Enable to use the header row in each sheet to define column names. Otherwise, the gem assigns generic column names automatically. | # Macro Source: https://docs.prophecy.ai/data-analysis/gems/custom/macro Use dbt macros in your pipelines This gem runs in . ## Overview The Macro gem lets you use a macro that you have defined or imported in your SQL project. Macros provide a simple interface where you can define the values of your macro parameters (arguments). Use the Macro gem when you: * Import macros via DBT Hub dependency. * Want to use a simple interface for custom gems. ## Parameters | Parameter | Description | | ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Macro | The macro that you wish to use.
Note: You can only select a Gem macro. You cannot select a macro defined as a [function](/data-analysis/development/extensibility/functions). | | Parameter `1` | A parameter that is defined in the macro. This is a value that is passed to the macro. | | Parameter `N` | Additional parameters that are defined in the macro. | ## Example: dbt\_utils For this example, assume you want to use the `dbt_utils` package as a [dependency](/data-analysis/development/extensibility/dependencies) in your project. 1. Open the **Options** (ellipses) menu in the project header. 2. Select **Dependencies**. 3. Click **+ Add Dependency**. 4. Choose the dependency type (Project, GitHub, or DBT Hub). 5. Add the package `dbt-labs/dbt_utils`. 6. Choose version `1.3.0`. 7. Click **Create**. 8. Click **Reload and Save** to download the package. 9. Review the functions and gems that are available in the dbt\_utils package. You can click any function or gem to see the underlying macro code. To use one of the imported gems in your pipeline: 1. Add the Macro gem to the canvas. 2. Open the gem configuration. 3. In the **Macro** field, choose the gem that you want to use. The gem parameters should appear. 4. Fill in the parameters. These are the arguments of the dbt macro. 5. Save and run the gem. The following image shows the unpivot macro imported from `dbt_utils` in a Macro gem. Unpivot macro # RestAPI Source: https://docs.prophecy.ai/data-analysis/gems/custom/rest-api Call REST APIs from your pipeline This gem runs in . ## Overview Use the RestAPI gem to make HTTP requests to external REST APIs from your pipeline. You can retrieve data, send data to downstream systems, or trigger external services as part of your workflow. ## Use cases Common uses for the RestAPI gem include: * Enriching pipeline data with information from an external service. * Sending pipeline results to another application. * Triggering external workflows or notifications. * Retrieving reference or lookup data during pipeline execution. ## Input and output The RestAPI gem uses the following input and output ports: | Port | Description | | ------- | ---------------------------------------------------------------------------------------- | | **in0** | (Optional) Input table. If connected, Prophecy makes one API request for each input row. | | **out** | Output table containing the API response, along with the input columns when applicable. | By default, the RestAPI gem does not include an input port. You can add one by clicking the **+** button next to **Ports**, or by connecting another gem directly to the RestAPI gem. You can change the number of input and output ports for most gems. To learn more, see [Gem ports](/data-analysis/gems/gems#gem-ports). ## Configure the RestAPI gem When you first open the RestAPI gem, Prophecy guides you through a multi-step configuration flow. The **Request** page defines the HTTP request that will be sent. Configure the following settings: | Setting | Description | | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------- | | HTTP Method | Select the HTTP method to use. Supported methods include `GET`, `POST`, `PUT`, `DELETE`, and `PATCH`. | | Target URL | Enter the URL of the REST API endpoint to call. | | Request Body | For methods that send a request body, choose the body format and provide the request payload. Supported formats include JSON, XML, and plain text. | You can reference input columns anywhere the gem accepts text by enclosing the column name in double curly braces: ```text theme={null} https://api.company.com/customers/{{CustomerID}} ``` You can also reference columns inside a request body: ```json theme={null} { "customerId": "{{CustomerID}}", "region": "{{Region}}" } ``` Once you have configured the HTTP request, click **Continue** to proceed. The **Auth & Headers** page lets you configure authentication and additional request metadata. Choose one of the supported authentication methods: * **None** β€” Send the request without automatic authentication. You can add an `Authorization` header manually in the **Request headers** section. * **Basic Auth** β€” Authenticate using a username and password stored in a [Username & Password secret](/data-analysis/environment/secrets/secrets). * **Bearer Token** β€” Authenticate using a bearer token. Select an existing secret to supply the token value, or enter a new token as a secret. * **API Key** β€” Send an API key in either a request header or query parameter. You can also configure: * **Request parameters** to append query parameters to the URL. For example, `limit:250` appends `?limit=250` to the request URL. * **Request headers** to send additional HTTP headers. Header values support hard-coded strings, pipeline parameters, and secrets. Store credentials as [Prophecy secrets](/data-analysis/environment/secrets/secrets) rather than hard-coding them in headers or parameters. Secrets are never visible after they are saved. Click **Continue** to proceed. The **Response** page controls how Prophecy parses the API response. Choose one of the following parsing options. ### None Return the API response as a single text column. You can specify a name for the output column. ### JSON Parse the response as JSON. You can optionally: 1. **Flatten the parsed response into columns.** If you leave this option disabled, Prophecy returns the parsed response as a single struct column. 2. **Wait for async results.** Enable this option for APIs that initially return an "in progress" response. Enter a JSONPath condition that identifies an in-progress response. Prophecy continues polling while the condition evaluates to true and stops when it evaluates to false or no longer matches. ### XML Parse the response as XML. You can optionally: 1. **Flatten the parsed response into columns.** If you leave this option disabled, Prophecy returns the parsed response as a single struct column. 2. **Wait for async results.** Enable this option for APIs that initially return an "in progress" response. Enter an XPath condition that identifies an in-progress response. Prophecy continues polling while the condition evaluates to true and stops when it evaluates to false or no longer matches. ### Parsed response output When you choose JSON or XML, the output depends on whether you flatten the parsed response. Parsing the response does not automatically define the output schema. After configuring the gem, open the output port and click **Infer from Cluster** to retrieve the response schema. | Flatten parsed response into columns | Output | | ------------------------------------ | -------------------------------------------------------------------------- | | Disabled | The parsed response is returned as a single struct column. | | Enabled | Fields from the parsed response are returned as individual output columns. | Click **Continue** to proceed. The **Resilience** page lets you configure retry behavior for failed requests. Enable **Retry if the API fails** to retry unsuccessful requests automatically. When retries are enabled, configure: * **Retry attempts** β€” Maximum number of retry attempts. * **Retry delay** β€” Number of seconds to wait between attempts. Prophecy retries requests that fail because of: * Network errors or timeouts. * HTTP `408 Request Timeout` responses. * HTTP `429 Too Many Requests` responses. * HTTP `5xx` server errors. Other HTTP `4xx` responses are not retried. Click **Save Configuration** to finish configuring the gem. ## Example Suppose you have a dataset of customer records with the following columns: | Column name | Description | | ------------- | ----------------------------------------------- | | `customer_id` | Unique identifier for each customer. | | `region` | Geographic region associated with the customer. | You can use the RestAPI gem to enrich each row by calling an external CRM API that returns account details for a given customer. ### Example configuration | Setting | Value | | ---------------- | ---------------------------------------------- | | HTTP Method | `GET` | | Target URL | `https://api.crm.com/accounts/{{customer_id}}` | | Authentication | Bearer Token | | Response parsing | JSON, flattened into columns | ### Result This configuration: * Sends one `GET` request per input row, using the `customer_id` column to construct the URL dynamically. * Authenticates each request using a bearer token stored as a Prophecy secret. * Parses the JSON response and flattens the returned fields into individual output columns alongside the original input columns. ## Common issues ### Authentication errors If requests return `401 Unauthorized` or `403 Forbidden`: * Verify that the token or credentials are correct and have not expired. * Check that the secret is attached to the correct fabric. * If using **None** authentication, confirm that the `Authorization` header is formatted correctly. Bearer tokens should use the format `Bearer `. ### Unexpected response shape If the output schema looks wrong after inferring: * Check the raw response by temporarily disabling JSON or XML parsing and inspecting the `api_data` column. * Some APIs wrap results in a top-level key such as `data` or `results`. You may need a subsequent gem to extract the nested field. ### Pagination Most APIs return a limited number of results per request, controlled by parameters such as `limit` and `offset` or `page`. The RestAPI gem sends one request per input row and does not handle pagination automatically. To retrieve all pages, build a loop using pipeline parameters or pre-compute the required page offsets in an upstream gem. ### API not reachable from Prophecy If the API call fails with a network or connection error, the API may not be accessible from Prophecy's execution environment. In this case, consider calling the API externally and writing the results to a table your fabric can read (such as a Databricks notebook writing to a Delta table, a Snowflake task, or a BigQuery scheduled query) then reading from that table as a source in your Prophecy pipeline. ## Similar tools and concepts | Tool or platform | Similar concept | | ---------------- | -------------------------------------------- | | SQL | `JOIN` against an external data source | | Alteryx | Download tool | | Pandas | `requests` library with `DataFrame.apply()` | | PySpark | UDF wrapping `requests` library | | Postman | Manual API testing and request configuration | # Script gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/custom/script Embed a custom Python script in your pipeline The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Where the script runs The execution environment for the Script gem depends on the [SQL warehouse](/data-analysis/environment/fabrics/prophecy-fabrics) configured in your Prophecy fabric. | SQL warehouse provider | Execution environment | | ---------------------- | --------------------- | | Databricks | Databricks Serverless | | Prophecy In Memory | Prophecy Automate | | BigQuery | Prophecy Automate | | Snowflake | Prophecy Automate | This ensures your Python logic runs in an environment optimized for your data platform. ## Parameters | Parameter | Description | | --------- | --------------------------------------- | | Script | Where you will write your Python script | Number of inputs and outputs can be changed as needed by clicking the `+` button on the respective tab. ## Troubleshooting ### Permission error with service principals When using a fabric with [service principal authentication for Databricks](/data-analysis/environment/connections/databricks), you may encounter the following error: ``` Failed due to: Unable to get run status for job id: INTERNAL_ERROR Cannot read the python file dbfs:/prophecy_tmp/prophecy_script_gem_[uuid].py User does not have permission SELECT on ANY File ``` This occurs because the Script gem uploads Python scripts to DBFS for execution. Service principals require explicit permissions to read files from DBFS. To solve this, contact your Databricks administrator for assistance. One workaround is to grant the service principal `SELECT` permission on `ANY FILE`: ```sql theme={null} GRANT SELECT ON ANY FILE TO ; ``` Granting `SELECT ON ANY FILE` provides broad read access to the file system. Consider the security implications for your environment before implementing this workaround. # SQLStatement gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/custom/sql-statement Use a custom SQL statement This gem runs in . ## Overview Use a custom SQL statement in your pipeline. The SQLStatement gem supports SELECT statements. This gem does not support actions like inserting or deleting tables. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Input and Output The SQLStatement gem uses the following input and output ports. | Port | Description | | ------- | ------------------------------------- | | **in0** | (Optional) Input table used in query. | | **out** | Output table with query results. | By default, the SQL statement gem does not include an input port. To add a port, click the `+` button next to **Ports**. Alternatively, connect the previous gem to the SQLStatement gem directly to add a new input port. Number of inputs and outputs can be changed as needed by clicking the `+` button on the respective tab. To learn more about adding and removing ports, see [Gem ports](/data-analysis/gems/gems#gem-ports). ## Parameters Configure the SQLStatement gem using the following parameters. | Parameter | Meaning | | --------- | ----------------------------------------- | | Out | SQL query that defines the output result. | Write your SQL query in the syntax of your SQL warehouse provider. For example, if you're connected to a Databricks SQL warehouse, use Databricks SQL dialect. If you're using the [Prophecy In Memory](/data-analysis/environment/fabrics/prophecy-fabrics), use DuckDB dialect. To reference the input table in the SQL query, use the name of the gem that generates the input table. ## Example Assume you have a gem named `weather_predictions`. It outputs the following table.
| DatePrediction | TemperatureCelsius | HumidityPercent | WindSpeed | Condition | | -------------- | ------------------ | --------------- | --------- | --------- | | 2025-03-01 | 15 | 65 | 10 | Sunny | | 2025-03-02 | 17 | 70 | 12 | Cloudy | | 2025-03-03 | 16 | 68 | 11 | Rainy | | 2025-03-04 | 14 | 72 | 9 | Sunny |
To filter out predictions before March 3: 1. Add a SQLStatement gem to the canvas. 2. Connect the `weather_predictions` gem to the SQLStatement gem directly in the canvas. Alternatively, add an input port in the gem configuration and choose the `weather_predictions` gem for **in0**. 3. In the code editor, paste the following query. ```sql theme={null} SELECT * FROM weather_predictions WHERE DatePrediction > '2025-03-02' ``` This query uses the gem name `weather_predictions` as the table name in the SELECT statement. 4. Run the gem. The following table appears as the gem output.
| DatePrediction | TemperatureCelsius | HumidityPercent | WindSpeed | Condition | | -------------- | ------------------ | --------------- | --------- | --------- | | 2025-03-03 | 16 | 68 | 11 | Rainy | | 2025-03-04 | 14 | 72 | 9 | Sunny |
# StoredProcedure Source: https://docs.prophecy.ai/data-analysis/gems/custom/stored-procedure Create and call stored procedures to use in pipelines This gem runs in . ## Overview Use the StoredProcedure gem to call stored procedures defined in your project. Stored procedures allow you to run procedural logic (such as loops, conditional statements, or DDL operations) as a step in your pipeline. Stored procedures are built on top of BigQuery stored procedures and follow the same SQL syntax and execution model. To learn more about what stored procedures are, how they work in Prophecy, and how to define them, see [Stored procedures](/data-analysis/gems/custom/stored-procedure). ## Prerequisites * Run Prophecy version 4.1.2 or higher. ## Limitations The StoredProcedure gem can only call stored procedures that have been defined in Prophecy. Stored procedures originating from BigQuery cannot be called from this gem. ## Input and Output By default, the StoredProcedure gem has no input or output ports. You can add ports manually by clicking `+` next to **Ports** in the gem configuration. * Input port (optional): Pass values from a table into the stored procedure. Only one input port is supported. * Output port (optional): Return output arguments in a table. Only one output port is supported. If an input port is added, the stored procedure is executed **once per input row**. As a result, the output will contain the **same number of rows** as the input. If no input port is added, the stored procedure runs once. ## Parameters Configure the StoredProcedure gem using the following parameters: | Parameter | Description | | -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Procedure | Select a stored procedure from the dropdown menu of [existing procedures](/data-analysis/gems/custom/stored-procedure) in the project. | | Arguments | Add values to arguments required by the stored procedure. Arguments automatically appear when you choose a stored procedure. | | Pass through columns | Pass through additional columns from the input table to the output of the stored procedure. Columns can be defined using visual or SQL expressions.
If no pass through columns are defined, the output contains one column per `OUT` argument in the stored procedure. | # ToDo Source: https://docs.prophecy.ai/data-analysis/gems/custom/todo-gem Create a placeholder gem in your pipeline The ToDo gem lets you create a placeholder in your pipeline for a future gem. Prophecy may insert ToDo gems automatically during migration from another tool when certain logic can't be translated. [Reach out](mailto:support@prophecy.io) to learn more about migration. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Input and Output Though the ToDo gem does not pass any data, you can pre-configure the output schema of the gem for future reference. | Port | Description | | ------- | ------------------------------------------------------------------------------- | | **in0** | The intended input for a future gem. | | **out** | The intended output for a future gem, for which you can create a custom schema. | ## Parameters | Parameter | Description | | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | | Highlight message | A custom message you write. It appears in [Diagnostics](/data-analysis/development/studio/studio#footer) to describe why this placeholder exists. | | Error message | A system-generated message added during migration. It explains why a process couldn't be migrated so you have context for resolving it. | | Helper code/text | Notes, code snippets, or context you provide to guide the development of a functional gem later on. | # Supported data types Source: https://docs.prophecy.ai/data-analysis/gems/data-types Review the set of data types supported in pipelines When processing data in Prophecy, the available data types depend on which SQL warehouse your organization uses. This guide outlines which data types are supported in Prophecy for each platform. You'll work with data types when reading data sources and defining their schemas. Infer schema ## Data types The tables below list the data types Prophecy supports depending on the [SQL warehouse](/data-analysis/environment/fabrics/prophecy-fabrics) you use for processing. The data types you see in Prophecy correspond to [data types in Databricks SQL](https://docs.databricks.com/aws/en/sql/language-manual/sql-ref-datatypes). | Data type | Description | | --------- | ----------------------------------------------------------------------------------------- | | Array | Represents values comprising a sequence of elements. | | Bigint | Represents 8-byte signed integer numbers. | | Binary | Represents byte sequence values. Only partially supported in the Prophecy UI. | | Boolean | Represents true and false values. | | Date | Represents values comprising year, month, and day, without a timezone. | | Decimal | Represents numbers with maximum precision and fixed scale. | | Double | Represents 8-byte double-precision floating point numbers. | | Float | Represents 4-byte single-precision floating point numbers. | | Integer | Represents 4-byte signed integer numbers. | | Smallint | Represents 2-byte signed integer numbers. | | String | Represents character string values. | | Struct | Represents values with the structure described by a sequence of fields. | | Timestamp | Represents values comprising year, month, day, hour, minute, and second, with a timezone. | | Tinyint | Represents 1-byte signed integer numbers. | | Variant | Represents semi-structured data. Only partially supported in the Prophecy UI. | | Void | Represents the untyped NULL. Only partially supported in the Prophecy UI. | These data types you see in Prophecy correspond to the [data types in BigQuery](https://cloud.google.com/bigquery/docs/reference/standard-sql/data-types). | Data type | Description | | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | | Record | Represents a column with [nested data](https://cloud.google.com/bigquery/docs/nested-repeated).
SQL type name: **STRUCT** | | Array | Represents a list of non-array values of the same data type.
SQL type name: **ARRAY** | | BigNumeric | Represents a decimal value with precision of 76.76 digits.
SQL type name: **BIGNUMERIC**
SQL aliases: **BIGDECIMAL** | | Numeric | Represents a decimal value with precision of 38 digits.
SQL type name: **NUMERIC**
SQL aliases: **DECIMAL** | | Datetime | Represents a Gregorian date and a time, independent of time zone.
SQL type name: **DATETIME** | | Time | Represents a time of day, independent of a specific date and time zone.
SQL type name: **TIME** | | Date | Represents a Gregorian calendar date, independent of time zone.
SQL type name: **DATE** | | Timestamp | Represents an absolute point in time, with microsecond precision and no timezone context.
SQL type name: **TIMESTAMP** | | Boolean | Represents true and false values.
SQL type name: **BOOL**
SQL aliases: **BOOLEAN** | | Float | Represents an approximate double-precision numeric value.
SQL type name: **FLOAT64** | | Integer | Represents a 64-bit integer.
SQL type name: **INT64**
SQL aliases: **INT**, **SMALLINT**, **INTEGER**, **BIGINT**, **TINYINT**, **BYTEINT** | | Bytes | Represents variable-length binary data.
SQL type name: **BYTES** | | String | Represents variable-length character strings.
SQL type name: **STRING** | * **JSON:** Supported. Prophecy infers the JSON schema and displays nested fields in the UI. * **Geography:** Serialized and stored as strings. * **Interval:** Converted to a month-day-nano interval type for display and processing.
The data types you see in Prophecy correspond to the [data types in Snowflake](https://docs.snowflake.com/en/sql-reference-data-types). | Data type | Description | | --------- | ------------------------------------------------------------------------- | | Array | Represents semi-structured arrays of values. | | Binary | Represents variable-length binary data. | | Boolean | Represents true and false values. | | Date | Represents calendar dates without a time component. | | Decimal | Represents fixed-point numbers with a specified precision and scale. | | Float | Represents approximate numeric values. | | Integer | Represents whole numbers. | | Object | Represents semi-structured key-value data. | | String | Represents variable-length character strings. | | Time | Represents a time of day without a date. | | Timestamp | Represents a date and time value. | | Variant | Represents semi-structured data such as JSON, Avro, ORC, Parquet, or XML. | Prophecy maps Snowflake SQL types to the simplified data types shown above. For example: * **Integer** includes Snowflake integer types such as `NUMBER`, `DECIMAL`, `NUMERIC`, `INT`, `INTEGER`, `BIGINT`, `SMALLINT`, `TINYINT`, and `BYTEINT` when they represent whole numbers. * **Float** includes approximate numeric types such as `FLOAT`, `FLOAT4`, `FLOAT8`, `DOUBLE`, `DOUBLE PRECISION`, and `REAL`. * **String** includes character and text types such as `VARCHAR`, `CHAR`, `CHARACTER`, `STRING`, and `TEXT`. * **Timestamp** includes `TIMESTAMP_NTZ`, `TIMESTAMP_LTZ`, and `TIMESTAMP_TZ`.
# Gems for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/gems Building blocks of your data pipelines Gems are functional units in a pipeline that perform tasks such as reading, transforming, writing, or handling other data operations. Combine gems on the canvas to build a pipeline, connecting each gem's output to the next gem's input. Use the Prophecy Agent to help you add gems throughout your pipeline. ## Categories You can select gems from the following categories. Each category opens a dropdown menu when you click it. | Category | Description | | ----------------- | ------------------------------------------------------------------- | | **Source/Target** | Read and write data from various data providers. | | **Transform** | Modify, enrich, or reshape data during processing. | | **Prepare** | Clean, structure, and optimize data for analysis. | | **Join** | Merge, split, or link datasets. | | **Parse** | Parse individual columns that contain formats like XML and JSON. | | **Report** | Share results through channels such as email or Tableau. | | **Custom** | Enhance and extend Prophecy's functionality. | | **Spatial** | Work with geographic data, such as locations, distances, and areas. | ## Bookmark favorite gems You can bookmark frequently used gems so they're easiliy available from the gem drawer. To bookmark a gem, click the star icon next to the gem in the category's dropdown menu. Starred gems appear in the Favorites section of the toolbar, where they're displayed as icon-only shortcuts. From Favorites, you can: * Click a gem to add it directly to the canvas. * Drag a gem onto the canvas. * Click the star again to remove it from your favorites. ## Interactive gem examples To test a gem hands-on, you can try the **interactive example** of the gem. If you search for a gem in the project sidebar, you can open the associated example and run the preconfigured pipeline! Gem example ## Expressions Many gems include expression components where you can implement custom or complex logic. You can either [build expressions visually](/data-analysis/gems/visual-expression-builder/visual-expression-builder), or use SQL syntax to write expression in code. Prophecy will automatically convert visual expressions to code expressions, and vice versa. The SQL dialect for expressions depends on the SQL warehouse connection of your fabric. Project code must match the requirements of the execution environment. As a result, expressions can differ slightly between projects that run on Databricks, Snowflake, or Google BigQuery. ## Gem instance When you click on a gem from the gem drawer, Prophecy adds the gem to your pipeline. The callouts in the image below correspond to the numbered UI elements described in the table. Gem instance | Callout | UI element | Description | | :-----: | ------------- | ----------------------------------------------------------------------------------------------------------- | | 1 | Gem label | The name of this particular gem instance. It must be unique within a given pipeline. | | 2 | Gem type name | The type of gem. | | 3 | Input ports | One or more ports that accept connections from upstream gems. | | 4 | Output ports | One or more ports that connect to downstream gems. | | 5 | Gem phase | The [phase](#gem-phase) for this gem instance, which defines the order in which gem instances are executed. | | 6 | Open | The button that lets you open the gem configuration. | | 7 | Run button | A button that runs the pipeline up to and including the gem. | | 8 | Action menu | A menu that includes options to change the phase of the gem, add run conditions, delete the gem, and more. | | 9 | Warning | Indicator that the gem contains errors to be fixed. | If you select one or more gems, you can copy and paste them within the same pipeline or across pipelines. However, you cannot paste across projects that use different languages (for example, from SQL to Scala). ## Gem configuration Once a gem is on the canvas, you can configure its behavior. Available fields available depend on the gem type. See the documentation for each gem to learn about its specific parameters. ### Gem ports Each gem handles ports differently. Some require specific input and output ports. Others allow only one port or none at all. Refer to the documentation for each gem to understand its port requirements. To add a port: 1. Open the gem to access its configuration. 2. On the left, choose the **Input** or **Output** tab under **Ports**. 3. Click the `+` button next to **Ports**. If the `+` button doesn't appear, the gem does not support adding more ports. To remove ports from a gem: 1. Open the gem to access its configuration. 2. On the left, choose the **Input** or **Output** tab under **Ports**. 3. Click the pencil icon to enter **Edit** mode. 4. Hover over the port you want to remove and click the trash icon. 5. Click **Done** to save your changes and exit **Edit** mode. To hide ports from the gem configuration, you can collapse the ports panel using the collapse icon next to **Ports**. ### Visual and code view Most gems can be configured in the **visual** view or the **code** view. Use the visual expression builder to populate fields in the visual view. Prophecy will automatically convert visual expressions into SQL expressions. You can edit these SQL statements or write your own in the code view. ## Add a gem from a connection You can quickly connect a new gem by hovering over the connector between two gems to reveal a `+` button. Click it to open the gem palette and insert a new gem directly at that point in the flow. * Hover over a gem's input port to add a gem before it. If a connection already feeds that port, the new gem is spliced into it automatically * Hover over a gem's output port to add a gem after it. If the port has no downstream connections, the new gem attaches directly. If the port already feeds one or more gems, you choose how to insert it: * Insert into the flow β€” the new gem is placed between the existing gems, so data now passes through it (A β†’ New Gem β†’ B). * Add as a separate path β€” the new gem reads the same output as a new branch, and existing connections stay unchanged. options for adding gem to output The gem palette only shows gems compatible with the selected insertion point, grouped by category. Use the search field at the top of the palette to filter by name instead of scrolling through categories. Prophecy adjusts the canvas layout automatically to fit the new gem, and the entire change β€” the new gem, any rewired connections, and any repositioned gems β€” undoes as a single step. ## Action menu Selecting one or more gems reveals a toolbar at the bottom of the canvas with the following buttons: | **Button** | **Description** | | ------------- | ------------------------------------------------------------------------- | | **Add to AI** | Adds [mention](data-analysis/ai/agent/chat/mentions) for the gem to Chat. | | **Explain** | Copilot provides an explanation of what the gem does in the pipeline. | | **Label** | Copilot renames the gem. | | **Actions** | Opens a dropdown with additional options, listed below. | | **Delete** | Remove the gem from the pipeline. | The **Actions** dropdown includes: | **Action** | **Description** | | -------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Add to AI** | Adds [mention](/data-analysis/ai/agent/chat/mentions) for gem to Chat. | | **Explain** | Copilot provides an explanation of what the gem does in the pipeline. | | **Fix** | Copilot resolves an error in the gem configuration. | | **Add Description** | Manually write a comment that appears as a tooltip above the gem. | | **Attach Annotation** | Lets you add an expanded label for the gem on the canvas. | | **Data Preview** | Enables the [Data Preview](/data-analysis/development/runs/data-explorer/data-explorer) icon on the gem's output so you can inspect sample data. When disabled, the icon remains unavailable. | | **Record Count** | Displays the number of rows produced by the gem next to its output. | | **Change Phase** | Change the [phase](#gem-phase) of the gem. | | **Group** | Combine the selected gems into a [container](/data-analysis/development/studio/containers). | | **Make Outgoing Connections Wireless** | Convert wired connections to wireless ones. See [Wireless connections](data-analysis/development/studio/wireless-connections) for more detail. | ### Applying actions to multiple gems You can select multiple gems and apply **Change Phase**, **Attach Annotation**, **Data Preview**, **Record Count**, and **Group** to all of them at once from the **Actions** dropdown. Bulk **Data Preview** and **Record Count** toggles skip gems that don't support them, such as canvas comments, target-process gems, and containers. If your selection has a mix of current values (for example, some selected gems have Data Preview on and others off), the checkbox shows an indeterminate state until you click it. To apply actions to multiple gems: 1. Drag to select multiple gems on the canvas. 2. Click **Actions** at the bottom of the canvas. 3. Select desired action. ### Gem phase In a data pipeline, the **phase** of a gem determines the sequence in which it runs. Here's how it works: * Gems are assigned a numerical phase (e.g., `0`, `1`, `-1`), where lower values run first. For example, a gem with phase `0` will execute before a gem with phase `1`. * When a gem runs, all its upstream gems must also run. This means that if a downstream gem has phase `0` and an upstream gem has phase `1`, the upstream gem will be grouped into phase `0` to ensure proper execution. * Because of this dependency, the phase assigned to the last gem in a branch determines the phase of the entire branch. This means that when configuring gem phases, you only need to focus on the *leaf nodes*β€”the final gems in each branch of the pipeline. Normally, pipeline branches run in parallel. Using gem phases, you can develop your pipeline to run in different stages. This can be useful when one part of the pipeline depends on the results of another, allowing you to control the execution order and ensure that data flows correctly from one stage to the next. # Except Source: https://docs.prophecy.ai/data-analysis/gems/join-split/except Return rows from the first dataset that do not appear in any of the others This gem runs in . ## Overview Use the Except gem to extract rows that are present in the **first table** but **absent** from all subsequent tables. This is useful for identifying gaps, such as missing orders, unprocessed records, or customers who haven't returned. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Input and Output | Port | Description | | ------- | --------------------------------------------------------------------------------------------- | | **in0** | The primary input table. | | **in1** | The second input table. Any matching rows from `in0` will be excluded from the output. | | **inN** | Optional: Additional input tables. Matching rows from `in0` will be excluded from the output. | | **out** | A table containing rows from `in0` that do **not** appear in any other input. | To add additional input ports, click `+` next to **Ports**. All input tables must have **identical schemas** (matching column names and data types). ## Parameters | Parameter | Description | | ----------------------- | ----------------------------------------------- | | Operation Type | Shows that the set operation type is `Except` | | Preserve duplicate rows | Checkbox to keep duplicates in the output table | ## Example Let's say you're working with two tables: **Table A** and **Table B**. * Both tables contain order-related data. * **Table A** contains order information from customer `1`, `2`, and `3`. * **Table B** contains order information from customer `1`, `2`, `3`, and `4`. * These tables contain some identical records (duplicates). ### Table A
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 101 | 1 | 2024-12-01 | 250.00 | | 102 | 2 | 2024-12-03 | 150.00 | | 101 | 1 | 2024-12-01 | 250.00 | | 104 | 3 | 2025-02-10 | 200.00 |
### Table B
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 103 | 1 | 2025-01-15 | 300.00 | | 104 | 3 | 2025-02-10 | 200.00 | | 105 | 4 | 2025-03-05 | 400.00 | | 106 | 2 | 2025-03-07 | 180.00 |
### Result #### Default The table that results from the Except gem only includes records in Table A that are not in Table B.
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 101 | 1 | 2024-12-01 | 250.00 | | 102 | 2 | 2024-12-03 | 150.00 |
The output indicates that order `101` and `102` appear in **Table A**, but not in **Table B**. #### Preserve duplicates If you select the `preserve duplicates` option, the gem preserves the duplicates from Table A.
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 102 | 2 | 2024-12-03 | 150.00 | | 101 | 1 | 2024-12-01 | 250.00 | | 101 | 1 | 2024-12-01 | 250.00 | The output is the same as above, except the gem also retains the duplicate order `101`.
# Intersect Source: https://docs.prophecy.ai/data-analysis/gems/join-split/intersect Return only the rows that are common across all input datasets This gem runs in . ## Overview Use the Intersect gem to return only the rows that appear in **all** input tables. This is useful for identifying overlapping data, such as customers who are active on multiple platforms, or transactions that appear across different systems or logs. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Input and Output | Port | Description | | ------- | ------------------------------------------------------------------------ | | **in0** | The first input table to compare. | | **in1** | The second input table to compare. | | **inN** | Optional: Additional tables to compare. | | **out** | A single table containing only rows that appear in **all** input tables. | To add additional input ports, click `+` next to **Ports**. All input tables must have **identical schemas** (matching column names and data types). ## Parameters | Parameter | Description | | ----------------------- | ------------------------------------------------ | | Operation Type | Shows that the set operation type is `Intersect` | | Preserve duplicate rows | Checkbox to keep duplicates in the output table | ## Example Let's say you're working with two tables: **Table A** and **Table B**. * Both tables contain order-related data. * **Table A** contains order information from customer `1`, `2`, and `3`. * **Table B** contains order information from customer `1`, `2`, `3`, and `4`. * **Tables A and B** each contain some identical records (duplicates). ### Table A
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 101 | 1 | 2024-12-01 | 250.00 | | 102 | 2 | 2024-12-03 | 150.00 | | 101 | 1 | 2024-12-01 | 250.00 | | 104 | 3 | 2025-02-10 | 200.00 |
### Table B
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 101 | 1 | 2024-12-01 | 250.00 | | 104 | 3 | 2025-02-10 | 200.00 | | 105 | 4 | 2025-03-05 | 400.00 | | 106 | 2 | 2025-03-07 | 180.00 | | 101 | 1 | 2024-12-01 | 250.00 |
### Result #### Default The table that results from the Intersect gem only includes the records present in all the input tables.
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 101 | 1 | 2024-12-01 | 250.00 | | 104 | 3 | 2025-02-10 | 200.00 |
The output indicates that order `101` and `104` appear in both **Table A** and **Table B**. By default, the duplicate row for order `101` is removed from the output. #### Preserve duplicates If the `preserve duplicates` option is selected, the duplicates present in **all** inputs are preserved.
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 101 | 1 | 2024-12-01 | 250.00 | | 101 | 1 | 2024-12-01 | 250.00 | | 104 | 3 | 2025-02-10 | 200.00 |
The output is the same as above, except the duplicate order `101` is retained. # Join gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/join-split/join Join two or more datasets This gem runs in . ## Overview Use the Join gem to combine related data from two or more datasets. Common use cases include: * adding customer details to order records * combining user activity with account information * enriching datasets with lookup tables * matching records between systems The Join gem matches rows using one or more join conditions, such as matching customer IDs or order numbers. By default, the Join gem provides a guided interface for choosing a join type, defining match conditions, and selecting output columns. If you need more than two input datasets or a condition that Simple mode can't represent, you can switch to Advanced mode to edit the underlying SQL configuration directly. ## Modes The Join gem supports two editing modes: | Mode | Description | | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | **Simple mode** (default) | Build a join using guided controls for the join type, match conditions, and output columns. Supports two input datasets. | | **Advanced mode** | Edit the full join configuration directly. Required for three or more input datasets, or conditions that Simple mode can't represent. | ## Simple mode Simple mode organizes the Join gem into three parts. | Section | Description | | ------------------ | ------------------------------------------------------------------------------------------------------ | | **Join type** | A dropdown at the top of the dialog. See [How to choose a join type](#how-to-choose-a-join-type). | | **Match rows on** | One row per condition, each comparing a column from the first input to a column from the second input. | | **Output columns** | Choose which columns from either input appear in the result, and optionally rename them. | The dialog also shows a **Where rows go** panel that explains which output port each row of your data lands on for the selected join type. See [Output ports in Simple mode](#output-ports-in-simple-mode). ### Match rows on Each row compares one column from the first input to one column from the second input using an operator: | Operator | Meaning | | -------- | --------------------- | | `=` | equals | | `!=` | not equals | | `<` | less than | | `<=` | less than or equal | | `>` | greater than | | `>=` | greater than or equal | Multiple rows are combined with `AND`. If your join needs `OR` logic, a function call, or any comparison Simple mode can't show as a row, the gem falls back to Advanced mode β€” see [Complex joins](#complex-joins). ### Output columns Pick any column from either input to include in the result. If you don't rename a column, Prophecy uses the column name as the output name automatically. Semi and anti joins only offer columns from the first input. A semi join keeps rows from the first input that have a match in the second, and an anti join keeps rows from the first input that don't. Neither one adds columns from the second input to the result. ## How to choose a join type | Join type | When to use it | | ------------------- | ---------------------------------------------------------------------------------------------------- | | **Inner Join** | Keep only matching rows. | | **Left Join** | Keep all rows from the first dataset. | | **Right Join** | Keep all rows from the second dataset. | | **Full Outer Join** | Keep all rows from both datasets. | | **Cross Join** | Combine every row from both datasets. | | **Semi Join** | Keep rows from the first dataset that have a match in the second, without adding any of its columns. | | **Anti Join** | Keep rows from the first dataset that have *no* match in the second. | Simple mode's join-type dropdown offers Inner, Left Outer, and Right Outer Join. Full Outer, Cross, Natural, and Semi/Anti joins are only available in Advanced mode. ## Complex joins Simple mode can represent a join when all of the following are true: * Exactly two input datasets are connected. * There's exactly one join condition connecting them. * The join type is Inner, Left Outer, or Right Outer (or another type already configured in Advanced mode that Simple mode recognizes). * The condition is one or more column comparisons combined with `AND` (no `OR`, no functions). * Every output column is a plain column pick, not a calculated expression like `concat(a, b)`. If any of these aren't true, Simple mode shows a message that the configuration is too complex to edit there. From this message you have two options: * **Switch to Advanced Mode** lets you edit the existing configuration directly. * **Reset and use Default Mode** clears the match conditions and output columns so you can start over with a simple inner join. **Reset and use Default Mode** permanently removes the existing match conditions and output columns. If the gem has more than two input datasets, only **Switch to Advanced Mode** is available β€” reducing the join to two inputs isn't something Reset can do for you. ## Common join examples | Goal | Recommended join type | | :------------------------------------------ | :-------------------- | | Keep only matching records from both tables | Inner Join | | Keep all records from the left table | Left Join | | Keep all records from both tables | Full Outer Join | | Combine every row from both tables | Cross Join | ## Input and Output | Port | Description | | ------- | ---------------------------------------------- | | **in0** | The first input table in the join. | | **in1** | The second input table in the join. | | **inN** | Optional: Additional input table for the join. | Adding a third input, or opening a gem that already has one, moves the Join gem to Advanced mode β€” Simple mode only represents two-input joins. ### Output ports in Simple mode In Simple mode the Join gem always has three output ports, in this order: | Port | Contents | | -------- | ---------------------------------------------------------------------- | | **out0** | Rows from the first input that didn't make it into the joined result. | | **out1** | The joined (matched) rows. | | **out2** | Rows from the second input that didn't make it into the joined result. | Which reject ports actually carry rows depends on the join type: | Join type | out0 (left reject) | out1 (matched) | out2 (right reject) | | ----------- | :----------------: | :------------: | :-----------------: | | Inner | rows | rows | rows | | Left Outer | empty | rows | rows | | Right Outer | rows | rows | empty | A Left Outer join, for example, keeps every row from the first input in the matched output, so `out0` never has anything to emit β€” but a right-side row with no match is dropped from the joined result, and shows up on `out2` instead of disappearing silently. Connect a reject port downstream if you want to inspect or route unmatched rows rather than losing them. ### Output port in Advanced mode | Port | Description | | ------- | ------------------------------------------------------- | | **out** | A single table that results from the join operation(s). | ## Advanced mode configuration Advanced mode exposes the underlying join configuration directly, instead of the guided Simple mode cards. ### Join conditions You can add one or more join conditions to the gem depending on the number of input tables added. Rows are matched using shared values, such as customer IDs, order IDs, or email addresses. | Parameters | Description | | -------------- | --------------------------------------------------------------------------------------------------------------------- | | Join type | The different join types you can choose from. These may vary by SQL provider. Learn about different join types below. | | Join condition | The condition that matches rows between tables. | If you want to use a type of join that is available in your SQL warehouse, you can type the name of that join directly in Prophecy. ### Expressions | Parameters | Description | | ----------- | ------------------------------------------------------------------------------------------------------------------------------------------- | | Expressions | Selects which columns appear in the output dataset. If left empty, Prophecy passes through all the input columns without any modifications. | ### Example Assume you have two tables: *orders* and *customers*. You want the orders table to include customer information, so you need to join the tables based on customer ID. You only want to preserve records in the output that have a match. To do so: 1. Connect **orders** to **in0** and **customers** to **in1**. 2. Choose **Inner Join** as the join type. 3. If using a visual expression, use the following join condition: **in0.CustomerID *equals* in1.customer\_id** 4. If using the code expression, use the following SQL join condition: **in0.CustomerID = in1.customer\_id** 5. Leave the **Expressions** tile empty. 6. Save and run the gem. ## Common issues ### Duplicate rows after joining Duplicate rows can occur when multiple rows in one table match the same row in another table. Verify that: * the join keys are unique when expected * the correct join type is selected * duplicate records do not exist in the input datasets ### Missing records in the output Rows may be excluded depending on the selected join type. For example: * Inner Join keeps only matching rows * Left Join keeps all rows from the left table * Full Outer Join keeps all rows from both tables In Simple mode, check the reject ports (`out0` / `out2`) β€” rows that don't appear in the matched output often show up there instead of being silently dropped. See [Output ports in Simple mode](#output-ports-in-simple-mode). ### Null values in joined columns Rows with `NULL` join keys may not match other rows. ### Ambiguous column names If both datasets contain columns with the same name, rename or qualify columns to avoid ambiguity. ## Join types Suppose there are two tables, *Employees* and *Departments*, with the following contents: ### Employees
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_ID | | :----------- | :------------- | :------------- | | 1 | Alice | 10 | | 2 | Bob | 20 | | 3 | Charlie | 30 | | 4 | David | NULL | | 5 | Eve | 20 |
### Departments
| DEPARTMENT\_ID | DEPARTMENT\_NAME | | :------------- | :--------------- | | 10 | HR | | 20 | Engineering | | 30 | Sales | | 40 | Marketing |
### INNER JOIN Inner Join will return columns from both the tables and only the matching records as long as the condition is satisfied. For example, if the Join condition provided was `employees.department_id = departments.department_id`, the sample query would be: ``` SELECT e.employee_id, e.employee_name, d.department_name FROM employees e INNER JOIN departments d ON e.department_id = d.department_id; ```
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_NAME | | :----------- | :------------- | :--------------- | | 1 | Alice | HR | | 2 | Bob | Engineering | | 5 | Eve | Engineering | | 3 | Charlie | Sales |
### LEFT JOIN / LEFT OUTER JOIN Left Join (or Left Outer join) will return columns from both the tables and match records with records from the left table. The result-set will contain null for the rows for which there is no matching row on the right side. For example, if the Join condition provided was `employees.department_id = departments.department_id`, the sample query would be: ``` SELECT e.employee_id, e.employee_name, d.department_name FROM employees e LEFT JOIN departments d ON e.department_id = d.department_id; ```
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_NAME | | :----------- | :------------- | :--------------- | | 1 | Alice | HR | | 2 | Bob | Engineering | | 3 | Charlie | Sales | | 4 | David | NULL | | 5 | Eve | Engineering |
### RIGHT JOIN / RIGHT OUTER JOIN Right Join (or Right Outer Join) returns matching rows from both tables and preserves all rows from the right table. Rows from the right table that do not have a match in the left table will contain `NULL` values for the left table columns. For example, if the join condition is `employees.department_id = departments.department_id`, the sample query would be: ```sql theme={null} SELECT e.employee_id, e.employee_name, d.department_name FROM employees e RIGHT OUTER JOIN departments d ON e.department_id = d.department_id; ```
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_NAME | | :----------- | :------------- | :--------------- | | 1 | Alice | HR | | 2 | Bob | Engineering | | 5 | Eve | Engineering | | 3 | Charlie | Sales | | NULL | NULL | Marketing |
### FULL JOIN / FULL OUTER JOIN Full Outer Join will return columns from both the tables and matching records with records from the left table and records from the right table. The result-set will contain NULL values for the rows for which there is no matching. For example, if the Join condition provided was `employees.department_id = departments.department_id`, the sample query would be: ``` SELECT e.employee_id, e.employee_name, d.department_name FROM employees e FULL OUTER JOIN departments d ON e.department_id = d.department_id; ```
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_NAME | | :----------- | :------------- | :--------------- | | 1 | Alice | HR | | 2 | Bob | Engineering | | 3 | Charlie | Sales | | 4 | David | NULL | | 5 | Eve | Engineering | | NULL | NULL | Marketing |
### CROSS JOIN Returns the Cartesian product of two datasets. It combines all rows from both tables. Cross Join will not have any Join conditions specified. For example, the sample query would be: ``` SELECT e.employee_id, e.employee_name, d.department_name FROM employees e CROSS JOIN departments d; ```
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_NAME | | :----------- | :------------- | :--------------- | | 1 | Alice | HR | | 1 | Alice | Engineering | | 1 | Alice | Sales | | 1 | Alice | Marketing | | 2 | Bob | HR | | 2 | Bob | Engineering | | 2 | Bob | Sales | | 2 | Bob | Marketing | | 3 | Charlie | HR | | 3 | Charlie | Engineering | | 3 | Charlie | Sales | | 3 | Charlie | Marketing | | 4 | David | HR | | 4 | David | Engineering | | 4 | David | Sales | | 4 | David | Marketing | | 5 | Eve | HR | | 5 | Eve | Engineering | | 5 | Eve | Sales | | 5 | Eve | Marketing |
### LEFT SEMI JOIN A semi join returns rows from the left table that have at least one matching row in the right table β€” but only the left table's columns come through. Unlike an Inner Join, a semi join never duplicates a left row when it has more than one match on the right. For example, if the join condition provided was `employees.department_id = departments.department_id`, the sample query would be: ``` SELECT e.employee_id, e.employee_name, e.department_id FROM employees e LEFT SEMI JOIN departments d ON e.department_id = d.department_id; ```
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_ID | | :----------- | :------------- | :------------- | | 1 | Alice | 10 | | 2 | Bob | 20 | | 3 | Charlie | 30 | | 5 | Eve | 20 |
### LEFT ANTI JOIN An anti join returns rows from the left table that have *no* matching row in the right table β€” the inverse of a semi join. Only the left table's columns come through. For example, if the join condition provided was `employees.department_id = departments.department_id`, the sample query would be: ``` SELECT e.employee_id, e.employee_name, e.department_id FROM employees e LEFT ANTI JOIN departments d ON e.department_id = d.department_id; ```
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_ID | | :----------- | :------------- | :------------- | | 4 | David | NULL |
### NATURAL INNER JOIN A natural join (or Natural Inner Join) is identical to an explicit Inner Join but it automatically joins columns with the same names in both tables. Natural Join will not have any join conditions specified. For example, the sample query would be: ``` SELECT e.employee_id, e.employee_name, d.department_name FROM employees e CROSS JOIN departments d; ```
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_NAME | | :----------- | :------------- | :--------------- | | 1 | Alice | HR | | 2 | Bob | Engineering | | 5 | Eve | Engineering | | 3 | Charlie | Sales |
### NATURAL LEFT OUTER JOIN A natural Left Outer join (or Natural Left Join) is identical to an explicit Left Outer Join but it automatically joins columns with the same names in both tables. Natural Left Outer Join will not have any join conditions specified. For example, the sample query would be: ``` SELECT e.employee_id, e.employee_name, d.department_name FROM employees e NATURAL LEFT OUTER JOIN departments d; ```
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_NAME | | :----------- | :------------- | :--------------- | | 1 | Alice | HR | | 2 | Bob | Engineering | | 3 | Charlie | Sales | | 4 | David | NULL | | 5 | Eve | Engineering |
### NATURAL RIGHT OUTER JOIN A natural Left Right join (or Natural Right Join) is identical to an explicit Right Outer Join but it automatically joins columns with the same names in both tables. Natural Right Outer Join will not have any join conditions specified. For example, the sample query would be: ``` SELECT e.employee_id, e.employee_name, d.department_name FROM employees e NATURAL RIGHT OUTER JOIN departments d; ```
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_NAME | | :----------- | :------------- | :--------------- | | 1 | Alice | HR | | 2 | Bob | Engineering | | 5 | Eve | Engineering | | 3 | Charlie | Sales | | NULL | NULL | Marketing |
### NATURAL FULL OUTER JOIN A natural Full Outer join (or Natural Full Join) is identical to an explicit Full Outer Join but it automatically joins columns with the same names in both tables. Natural Full Outer Join will not have any join conditions specified. For example, the sample query would be: ``` SELECT e.employee_id, e.employee_name, d.department_name FROM employees e NATURAL FULL OUTER JOIN departments d; ```
| EMPLOYEE\_ID | EMPLOYEE\_NAME | DEPARTMENT\_NAME | | :----------- | :------------- | :--------------- | | 1 | Alice | HR | | 2 | Bob | Engineering | | 3 | Charlie | Sales | | 4 | David | NULL | | 5 | Eve | Engineering | | NULL | NULL | Marketing |
## Similar tools and concepts The Join gem combines related data from multiple datasets using shared column values. You may recognize similar behavior from: * SQL `JOIN` operations * the Alteryx Join tool * PySpark `join()` * Pandas `merge()` # Union Source: https://docs.prophecy.ai/data-analysis/gems/join-split/union Perform addition of rows from multiple tables This gem runs in . ## Overview The Union gem lets you combine records from different tables for cases such as merging customer databases or aggregating logs from multiple servers. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Input and Output | Port | Description | | ------- | ----------------------------------------------------- | | **in0** | The first input table. | | **in1** | The second input table. | | **inN** | Optional: Additional input tables. | | **out** | A single table containing all rows from input tables. | To add additional input ports, click `+` next to **Ports**. All input tables must have **identical schemas** (matching column names and data types). If table columns have different orders or a different number of columns, use the [UnionByName](/data-analysis/gems/join-split/union-by-name) gem. ## Parameters | Parameter | Description | | ----------------------- | ----------------------------------------------- | | Operation Type | Shows that the set operation type is `Union` | | Preserve duplicate rows | Checkbox to keep duplicates in the output table | ## Example Let's say you're working with two tables: **Table A** and **Table B**. * Both tables contain order-related data. * **Table A** contains order information from customer `1`, `2`, and `3`. * **Table B** contains order information from customer `1`, `2`, `3`, and `4`. * These tables contain some identical records (duplicates). ### Table A
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 101 | 1 | 2024-12-01 | 250.00 | | 102 | 2 | 2024-12-03 | 150.00 | | 103 | 1 | 2025-01-15 | 300.00 | | 104 | 3 | 2025-02-10 | 200.00 |
### Table B
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 103 | 1 | 2025-01-15 | 300.00 | | 104 | 3 | 2025-02-10 | 200.00 | | 105 | 4 | 2025-03-05 | 400.00 | | 106 | 2 | 2025-03-07 | 180.00 |
### Result without duplicates The following is the output table without duplicates.
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 101 | 1 | 2024-12-01 | 250.00 | | 102 | 2 | 2024-12-03 | 150.00 | | 103 | 1 | 2025-01-15 | 300.00 | | 104 | 3 | 2025-02-10 | 200.00 | | 105 | 4 | 2025-03-05 | 400.00 | | 106 | 2 | 2025-03-07 | 180.00 |
### Result with duplicates The following is the output table when you select **Preserve duplicate rows** in the gem configuration.
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 101 | 1 | 2024-12-01 | 250.00 | | 102 | 2 | 2024-12-03 | 150.00 | | 103 | 1 | 2025-01-15 | 300.00 | | 104 | 3 | 2025-02-10 | 200.00 | | 103 | 1 | 2025-01-15 | 300.00 | | 104 | 3 | 2025-02-10 | 200.00 | | 105 | 4 | 2025-03-05 | 400.00 | | 106 | 2 | 2025-03-07 | 180.00 |
Order `103` and `104` are duplicate records. # UnionByName Source: https://docs.prophecy.ai/data-analysis/gems/join-split/union-by-name Combine datasets by aligning columns with the same name This gem runs in . ## Overview Use the UnionByName gem to combine rows from multiple datasets by matching column names. This can help when working with data from different sources where schemas might vary slightly in order or structure, but the column names are consistent. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Input and Output | Port | Description | | ------- | -------------------------------------------------------------------------- | | **in0** | The first input table. | | **in1** | The second input table. | | **inN** | Optional: Additional input tables to include in the union. | | **out** | A single table containing the combined rows, with columns matched by name. | To add additional input ports, click `+` next to **Ports**. ## Parameters | Parameter | Description | | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Operation Type | Choose between two different operations.
  • Union By Name (No Missing Columns): All columns must be present in both tables.
  • Union By Name (Allow Missing Columns): Columns do not have to match between tables.
| ## Example Let's say you're working with two tables: **Table A** and **Table B**. * Both tables contain order-related data. * Table B is missing the `amount` column that exists in Table A. * The columns in **Table B** that match **Table A** are ordered differently. ### Table A
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 101 | 1 | 2024-12-01 | 250.00 | | 102 | 2 | 2024-12-03 | 150.00 | | 103 | 1 | 2025-01-15 | 300.00 | | 104 | 3 | 2025-02-10 | 200.00 |
### Table B
| `order_id` | `order_date` | `customer_id` | | ---------- | ------------ | ------------- | | 103 | 2025-01-15 | 1 | | 104 | 2025-02-10 | 3 | | 105 | 2025-03-05 | 4 | | 106 | 2025-03-07 | 2 |
### Result The table below shows the output after running a UnionByName gem with **Allow Missing Columns** enabled. Since **Table B** doesn't have the `amount` column, Prophecy fills in null values for rows coming from Table B.
| `order_id` | `customer_id` | `order_date` | `amount` | | ---------- | ------------- | ------------ | -------- | | 101 | 1 | 2024-12-01 | 250.00 | | 102 | 2 | 2024-12-03 | 150.00 | | 103 | 1 | 2025-01-15 | 300.00 | | 104 | 3 | 2025-02-10 | 200.00 | | 103 | 1 | 2025-01-15 | NULL | | 104 | 3 | 2025-02-10 | NULL | | 105 | 4 | 2025-03-05 | NULL | | 106 | 2 | 2025-03-07 | NULL |
This example will fail if you select **No Missing Columns** in the gem settings, since **Table B** is missing a column in **Table A**. # JSONParse Source: https://docs.prophecy.ai/data-analysis/gems/parse/json-parse Extract and parse JSON from a column This gem runs in . ## Overview Use the JSONParse gem to extract and parse JSON from a column within your table using configurable schema detection methods for structured, usable output. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Parameters | Parameter | Description | | ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Select column to parse | Specifies the input column containing the JSON data to be parsed. | | Parsing method | Determines how Prophecy derives the schema used to parse the JSON structure.
  • **Parse from sample record**. Prophecy uses the schema from the sample record you provide.
  • **Parse from schema**. Prophecy uses the schema you provide in the form of a schema struct.
| ## Output The output schema of the JSONParse gem includes all of the input columns and the parsed content as a **struct** data type. JSONParse Output # Regex Source: https://docs.prophecy.ai/data-analysis/gems/parse/regex Pattern matching and text extraction using regular expressions This gem runs in . ## Overview The Regex gem enables pattern matching and text extraction using regular expressions. This gem provides four distinct output methods for processing text data: * [Replace](#replace-configuration) * [Tokenize](#tokenize-configuration) * [Parse](#parse-configuration) * [Match](#match-configuration) ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Input and Output The Regex gem uses the following ports: | Port | Description | | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **in0** | The source table containing text data that needs to be processed with regex patterns. | | **out** | The output table containing:
  • Original columns preserved
  • New columns created based on the selected output method (Replace, Tokenize, Parse, or Match)
The output schema depends on the chosen method and configuration. | ## Parameters Configure the Regex gem using the following parameters. ### Common configuration These parameters are available for all regex operations: | Parameter | Description | | ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Select Column to Split | Choose the input column containing the text data you want to process with regex patterns. | | Output Method | Select how the regex operation should handle matches:
  • Replace: Substitute matched text with replacement values.
  • Tokenize: Split text into tokens or columns based on regex patterns.
  • Parse: Extract specific groups from regex matches into separate columns.
  • Match: Determine whether text matches the pattern.
| | Regex | Enter your regular expression pattern. The field supports standard regex syntax with capture groups for extracting specific portions of matched text.

**Note**: Different SQL dialects may require specific escaping. For example, to match a literal dot, generic regex uses `a\.` while Databricks SQL requires `a\\.` | | Case Insensitive Matching | Enable this option to perform pattern matching without regard to letter case. | ### Replace configuration The **Replace** method substitutes matched portions of text with specified replacement values. When using this method, the gem outputs an additional column with the replaced values. | Parameter | Description | | ----------------------------- | --------------------------------------------------------------------------- | | Replacement Text | Specify replacement text or use capture group references. | | Copy Unmatched Text to Output | When enabled, non-matching text is preserved in the appended output column. | #### Example Use this method to standardize phone number formats from `555-123-4567` to `(555) 123-4567`. * **Select Column to Split**: `phone_number` * **Regex**: `(\d{3})-(\d{3})-(\d{4})` * **Replacement text**: `($1)$2-$3` This inserts capture groups `1`, `2`, and `3` into the replacement pattern to create the new formatted string. The result is written to a new output column, while the original value is preserved. **Input table**
| id | phone\_number | | -- | ------------- | | 1 | 555-332-1234 | | 2 | 555-034-9876 |
**Output table**
| id | phone\_number | phone\_number\_replaced | | -- | ------------- | ----------------------- | | 1 | 555-332-1234 | (555)332-1234 | | 2 | 555-034-9876 | (555)034-9876 |
### Tokenize configuration The **Tokenize** method splits text into tokens based on regex patterns and capture groups. Each capture group becomes a token. This method creates either new columns or rows depending on your configuration. | Parameter | Description | | ------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Select Split Strategy | Choose how to split the data:
  • **Split to columns**: Breaks text into tokens and places each token into a new column in the same row. Requires a fixed number of columns.
  • **Split to rows**: Breaks text into tokens and outputs each token as a new row in a single column. Rows are generated dynamically, making this option useful when the number of tokens varies.
| | Allow Blank Tokens (Split to columns only) | If there are fewer tokens than the defined number of columns, allow empty strings to fill the extra columns. Otherwise, those columns are set to `NULL`. | | Number of columns (Split to columns only) | Specify the number of output columns to create for tokenized data. | | For Extra Columns (Split to columns only) | Define how to handle cases where there are more tokens than columns.
  • **Drop Extra with Warning**: Skip writing excess tokens and log a warning message to indicate this.
  • **Drop Extra without Warning**: Skip writing excess tokens silently without generating warnings.
  • **Error**: Stop processing and raise an error when the number of tokens exceeds the defined number of columns.
| | Output Root Name | Base name for the new column(s) containing the tokens. | #### Example Use this method to parse email addresses into username and domain components. * **Select Column to Split**: `email` * **Regex**: `([^@]+)@(.+)` * **Select Split Strategy**: Split to columns * **Number of columns**: 2 * **Output root name**: `token` **Input table**
| id | email | | -- | --------------------- | | 1 | `support@example.com` | | 2 | `sales@company.org` |
**Output table**
| id | email | token\_1 | token\_2 | | -- | --------------------- | --------- | ------------- | | 1 | `support@example.com` | `support` | `example.com` | | 2 | `sales@company.org` | `sales` | `company.org` |
### Parse configuration The **Parse** method extracts capture groups from regex matches and outputs each group as a separate column. Prophecy automatically generates one output column for every capture group in the regex. | Parameter | Description | | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | New Column Name | Specify the name for the new column. | | Select Data Type | Choose the data type. | | Regex Expression | View the capture group that will populate the column.
If you edit this value, Prophecy will automatically revert it to the original value. | Rows in the **Parse Configuration** table are determined by the number of capture groups in the **Regex** field. You cannot add additional rows to or remove rows from this table. #### Example Use this method to parse phone numbers into `area_code`, `exchange`, and `number` columns. * **Select Column to Split**: `phone_number` * **Regex**: `([0-9]{3})-([0-9]{3})-([0-9]{4})` * **Parse Configuration**: | New Column Name | Select Data Type | Regex Expression | | --------------- | ---------------- | ---------------- | | `area_code` | String | `([0-9]{3})` | | `exchange` | String | `([0-9]{3})` | | `number` | String | `([0-9]{4})` | **Input table**
| id | phone\_number | | -- | ------------- | | 1 | 555-332-1234 | | 2 | 555-034-9876 |
**Output table**
| id | phone\_number | area\_code | exchange | number | | -- | ------------- | ---------- | -------- | ------ | | 1 | 555-332-1234 | 555 | 332 | 1234 | | 2 | 555-034-9876 | 555 | 034 | 9876 |
### Match configuration The **Match** method determines whether text matches the specified regex pattern. Adds a column with 1 for matches and 0 for non-matches. | Parameter | Description | | ---------------------------- | ----------------------------------------------------------------------------------------------------------- | | Column name for match status | Specify the name for the new column containing match results. | | Error if not Matched | Enable to raise an error when no match is found. When disabled, non-matching rows will receive a `0` value. | #### Example Use this method to validate email addresses and create a binary match column. * **Select Column to Split**: `email` * **Regex**: `^[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}$` * **Column name for match status**: `is_valid_email` * **Error if not matched**: `Disabled` **Input table**
| id | email | | -- | ------------------------- | | 1 | `support@example.com` | | 2 | `sales.team` | | 3 | `engineering@company.org` |
**Output table**
| id | email | is\_valid\_email | | -- | ------------------------- | ---------------- | | 1 | `support@example.com` | 1 | | 2 | `sales.team` | 0 | | 3 | `engineering@company.org` | 1 |
# TextToColumns Source: https://docs.prophecy.ai/data-analysis/gems/parse/text-to-column Convert text into a column in your table This gem runs in . ## Overview When working with certain tables, you might encounter text columns that contain multiple values separated by specific characters such as commas or semicolons. Use the `TextToColumns` gem to parse this text and simplify further analysis and processing. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Parameters | Parameter | Description | | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Select Column to Split | The column that contains the text you would like to split. | | Delimiter | The character that delimits the separate values. | | Select Split Strategy |
  • Split to columns: Values will be split into individual rows in one or more columns.
  • Split to rows: Values will be split into individual rows in one column.
| ## Example Let's say you have the following table that includes bank account information. Note that some values in the `Beneficiaries` column contain multiple names separated by semicolons. | `Account_Number` | `Account_Type` | `Balance` | `Beneficiaries` | | ---------------- | -------------- | --------- | ---------------------------------------- | | 123456789 | Checking | 75000.50 | `Amina Yusuf; Juan PΓ©rez` | | 987654321 | Savings | 12000.00 | `Chen Wei` | | 456789123 | Checking | 3400.50 | `Yuki Tanaka; Leila Haddad; Ivan Petrov` | | 789123456 | Checking | 800.00 | `NULL` | | 321654987 | Savings | 2200.95 | `Mei Lin; Noah Schmidt` | You can use the `TextToColumns` gem to automatically split these values into separate rows or columns, making your data cleaner and easier to work with. 1. Open the `TextToColumns` gem. 2. Select the `Beneficiaries` column to split. 3. Under **Delimiter**, input `;` as the delimiting character. 4. Select the split strategy. Let's explore the outputs that result from the different split strategies. ### Split to columns When you use the **Split to columns** strategy, each name should be split into its own column. 1. Select **Split to columns**. 2. Under **Number of columns**, type `2`. 3. Keep the default **Extra Characters** setting to `Leave extra in last column`. 4. Keep or change the default column prefix and suffix. The resulting table will have two new columns. The output table has the same number of rows as the input table. | `Account_Number` | `Account_Type` | `Balance` | `Beneficiaries` | `root_1_generated` | `root_2_generated` | | ---------------- | -------------- | --------- | ---------------------------------------- | ------------------ | --------------------------- | | 123456789 | Checking | 75000.50 | `Amina Yusuf; Juan PΓ©rez` | `Amina Yusuf` | `Juan PΓ©rez` | | 987654321 | Savings | 12000.00 | `Chen Wei` | `Chen Wei` | `NULL` | | 456789123 | Checking | 3400.50 | `Yuki Tanaka; Leila Haddad; Ivan Petrov` | `Yuki Tanaka` | `Leila Haddad; Ivan Petrov` | | 789123456 | Checking | 800.00 | `NULL` | `NULL` | `NULL` | | 321654987 | Savings | 2200.95 | `Mei Lin; Noah Schmidt` | `Mei Lin` | `Noah Schmidt` | Notice that one the of cells still has a semicolon: `Leila Haddad; Ivan Petrov`. This is because the gem was configured to generate two new columns. Extra characters are kept in the last column. If the gem was configured to generate three columns instead, each beneficiary would have their own column. ### Split to rows The **Split to rows** strategy creates a separate row for each value in the selected column. All other column values are copied to each new row. Use this strategy when: * The number of items in a cell (such as beneficiaries) varies between rows. * You don't know the maximum number of items in the column. To apply this strategy: 1. Select **Split to rows**. 2. Keep or change the default generated column name. In the output table, each beneficiary has their own row. The output table has more rows than the input table. | `Account_Number` | `Account_Type` | `Balance` | `Beneficiaries` | `generated_column` | | ---------------- | -------------- | --------- | ---------------------------------------- | ------------------ | | 123456789 | Checking | 75000.5 | `Amina Yusuf; Juan PΓ©rez` | `Amina Yusuf` | | 123456789 | Checking | 75000.5 | `Amina Yusuf; Juan PΓ©rez` | `Juan PΓ©rez` | | 987654321 | Savings | 12000 | `Chen Wei` | `Chen Wei` | | 456789123 | Checking | 3400.5 | `Yuki Tanaka; Leila Haddad; Ivan Petrov` | `Yuki Tanaka` | | 456789123 | Checking | 3400.5 | `Yuki Tanaka; Leila Haddad; Ivan Petrov` | `Leila Haddad` | | 456789123 | Checking | 3400.5 | `Yuki Tanaka; Leila Haddad; Ivan Petrov` | `Ivan Petrov` | | 789123456 | Checking | 800 | `NULL` | `NULL` | | 321654987 | Savings | 2200.95 | `Mei Lin; Noah Schmidt` | `Mei Lin` | | 321654987 | Savings | 2200.95 | `Mei Lin; Noah Schmidt` | `Noah Schmidt` | # XMLParse Source: https://docs.prophecy.ai/data-analysis/gems/parse/xml-parse Parse XML inside a table This gem runs in . ## Overview Use the XMLParse gem to extract and convert XML data stored within a column of your table into a structured format. This gem lets you define how Prophecy parses the XML, either by using a sample record or a user-provided schema. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Parameters | Parameter | Description | | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Select column to parse | Specifies the input column containing the XML data to be parsed. | | Parsing method | Determines how Prophecy derives the schema used to parse the XML structure.
  • **Parse from sample record**. Prophecy uses the schema from the sample record you provide.
  • **Parse from schema**. Prophecy uses the schema you provide in the form of a schema struct.
| ## Output The output schema of the XMLParse gem includes all of the input columns and the parsed content as a **struct** data type. XMLParse # DataCleansing gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/prepare/data-cleansing Standardize data formats This gem runs in . ## Overview Use the DataCleansing gem to standardize data formats and address missing or null values in the data. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Parameters | Parameter | Description | | -------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | | Remove nulls from entire dataset | Removes any rows that contain null values.
This operates on all columnsβ€”not just those you selected to clean. | | Select columns to clean | Specifies the columns to apply data cleansing transformations to. | | Replace null values in column | Replaces null values in selected columns with a specified default.
Example: `0` for numeric columns, empty string for text | | Remove unwanted characters | Removes specified characters from all values in the selected columns.
Example: remove whitespaces or punctuation | | Modify case | Converts text in selected columns to a specified case format.
Example: lowercase, UPPERCASE, Title Case | ## Example Assume you have a dataset that includes all entries from a feedback survey.
| Name | Date | Rating | Feedback | | ----- | ---------- | ------ | -------------------------- | | Ada | 2025-04-18 | 5 | I really enjoy the product | | scott | 2025-04-18 | 5 | NULL | | emma | 2025-04-17 | 2 | The product is confusing | | NULL | 2025-04-17 | 3 | NULL |
The following is one way to configure a DataCleansing gem for this table: 1. Select columns to clean: `Name` 2. Replace null values in column: `Not provided` 3. Modify case: `Title Case` ### Result After the transformation, the table will look like:
| Name | Date | Rating | Feedback | | ------------ | ---------- | ------ | -------------------------- | | Ada | 2025-04-18 | 5 | I really enjoy the product | | Scott | 2025-04-18 | 5 | NULL | | Emma | 2025-04-17 | 2 | The product is confusing | | Not provided | 2025-04-17 | 3 | NULL |
# Deduplicate gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/prepare/deduplicate Remove duplicates from your data This gem runs in . ## Overview When working with data, it's common to run into duplicate information. Duplicates can come from multiple data sources, system errors, or repeated updates over time. Leverage the Deduplication gem to remove these duplicates, and pay close attention to the Deduplication mode that you use. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Parameters | Parameter | Description | | ------------------- | -------------------------------------------------------------- | | Mode | Deduplication method | | Expression | Column(s) to check for duplicates | | Use Custom Order By | Sort rows before deduplicating (First and Last mode **only**). | ## Mode Next to **Deduplicate On Columns**, choose how to keep certain rows. | Mode | Description | Output | | ------------- | --------------------------------------------------------------------------- | ------------------------------------------------------------------- | | Distinct Rows | Keeps one version of each duplicated row, removing extra duplicates. | All columns are passed through unless target columns are specified. | | Unique Only | Keeps only the rows that appear exactly once and removes any duplicate rows | All columns are passed through. | | First | Keeps the first occurrence of the duplicate row. | All columns are passed through. | | Last | Keeps the last occurrence of the duplicate row. | All columns are passed through. | ## Example Assume you have a table of contact information where some people appear more than once. This could be due to updates over time or repeated data entry. You want to identify and clean up these duplicates based on fields like email or phone number. Here is your original table:
| email | phone | first\_name | last\_name | date\_added | | ---------------------- | ------------ | ----------- | ---------- | ----------- | | `alex.t@example.com` | 123-456-7890 | Alex | Taylor | 2023-01-01 | | `alex.t@example.com` | 123-456-7890 | Alex | Taylor | 2023-07-01 | | `sam.p@example.com` | 987-654-3210 | Sam | Patel | 2024-03-15 | | `casey.l@example.com` | 555-111-2222 | Casey | Lee | 2024-05-01 | | `casey.l@example.com` | 555-111-2222 | Casey | Lee | 2025-01-01 | | `jordan.k@example.com` | 333-444-5555 | Jordan | Kelly | 2023-09-10 | | `morgan.s@example.com` | 666-777-8888 | Morgan | Smith | 2025-01-01 |
### Distinct Rows Let's look at what happens with our original table when using **Distinct Rows** without selecting any specific columns. If two rows match exactlyβ€”every value in every columnβ€”they are considered duplicates, and only one will be kept.
| email | phone | first\_name | last\_name | date\_added | | ---------------------- | ------------ | ----------- | ---------- | ----------- | | `alex.t@example.com` | 123-456-7890 | Alex | Taylor | 2023-01-01 | | `alex.t@example.com` | 123-456-7890 | Alex | Taylor | 2023-07-01 | | `sam.p@example.com` | 987-654-3210 | Sam | Patel | 2024-03-15 | | `casey.l@example.com` | 555-111-2222 | Casey | Lee | 2024-05-01 | | `casey.l@example.com` | 555-111-2222 | Casey | Lee | 2025-01-01 | | `jordan.k@example.com` | 333-444-5555 | Jordan | Kelly | 2023-09-10 | | `morgan.s@example.com` | 666-777-8888 | Morgan | Smith | 2025-01-01 |
The result is identical to the input because no two rows are exact matches across all columns. ### Distinct Rows with Target Columns Often, you want to identify duplicates based on just one or a few columns, like `email`. You can do this by selecting the columns to deduplicate on. Here is the result when email is selected as the target column:
| email | | ---------------------- | | `alex.t@example.com` | | `sam.p@example.com` | | `casey.l@example.com` | | `jordan.k@example.com` | | `morgan.s@example.com` |
Only the distinct values from the email column are returned. Other columns are not included, because Prophecy doesn't assume which corresponding values (like `phone` or `date_added`) to keep. If you want to keep related columns, use the First or Last deduplication methods. ### First and Last The **First** and **Last** options help you keep just one row for each duplicate based on the order of the data. You'll typically combine this with sorting to control which version is retained. For example, if you want to keep the most recent entry for each duplicate email address, you can: * Choose **First** and sort by `date_added` descending * Choose **Last** and sort by `date_added` ascending Here's the result when deduplicating by `email`, keeping the most recent record for each:
| email | phone | first\_name | last\_name | date\_added | | ---------------------- | ------------ | ----------- | ---------- | ----------- | | `alex.t@example.com` | 123-456-8888 | Alex | Taylor | 2023-07-01 | | `sam.p@example.com` | 987-654-3210 | Sam | Patel | 2024-03-15 | | `casey.l@example.com` | 555-111-2222 | Casey | Lee | 2025-01-01 | | `jordan.k@example.com` | 333-444-5555 | Jordan | Kelly | 2023-09-10 | | `morgan.s@example.com` | 666-777-8888 | Morgan | Smith | 2025-01-01 |
### Unique Only The **Unique Only** option removes all rows with duplicates. It keeps only the rows that appear exactly once, based on the selected columns. For example, if you choose to deduplicate based on `email`, any row that contains an email address that appears more than once will be removed entirely. Here's the result when using **Unique Only** on the original table with `email` as the target column:
| email | phone | first\_name | last\_name | date\_added | | ---------------------- | ------------ | ----------- | ---------- | ----------- | | `sam.p@example.com` | 987-654-3210 | Sam | Patel | 2024-03-15 | | `jordan.k@example.com` | 333-444-5555 | Jordan | Kelly | 2023-09-10 | | `morgan.s@example.com` | 666-777-8888 | Morgan | Smith | 2025-01-01 |
# Filter gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/prepare/filter Keep only rows that match a condition This gem runs in . ## Overview Use the Filter gem to keep rows that match a condition and remove all other rows from the pipeline. Common use cases include: * filtering active customers * removing null values * keeping recent records * filtering by date ranges * selecting high-value transactions The Filter gem evaluates each row individually and keeps only rows that match the filter condition. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Parameters | Parameter | Description | Required | | :--------------- | :-------------------------------------------------------- | :------- | | Model | Input dataset to filter. | True | | Filter Condition | Boolean expression used to determine which rows are kept. | True | ## Common filter examples | Goal | Filter condition | | :---------------------- | :-------------------------- | | Keep active customers | `Status = 'ACTIVE'` | | Keep high-value orders | `OrderAmount > 1000` | | Filter recent orders | `OrderDate >= '2025-01-01'` | | Remove null values | `CustomerId IS NOT NULL` | | Keep specific countries | `Country IN ('US', 'CA')` | ## Example Assume you have the following weather prediction table.
| DatePrediction | TemperatureCelsius | HumidityPercent | WindSpeed | Condition | | -------------- | ------------------ | --------------- | --------- | --------- | | 2025-03-01 | 15 | 65 | 10 | Sunny | | 2025-03-02 | 17 | 70 | 12 | Cloudy | | 2025-03-03 | 16 | 68 | 11 | Rainy | | 2025-03-04 | 14 | 72 | 9 | Sunny |
Using the following filter condition: ```text theme={null} DatePrediction > '2025-03-02' ``` returns:
| DatePrediction | TemperatureCelsius | HumidityPercent | WindSpeed | Condition | | -------------- | ------------------ | --------------- | --------- | --------- | | 2025-03-03 | 16 | 68 | 11 | Rainy | | 2025-03-04 | 14 | 72 | 9 | Sunny |
### Using pipeline parameters in filter conditions You can reference [pipeline parameters](/data-analysis/development/parameters/parameters) in filter conditions to make filtering dynamic at runtime. In **Visual mode**, select **Configuration Variables** from the expression builder to insert a parameter directly. In **Code mode**, use Jinja syntax: ``` {{ var('parameter_name') }} ``` | Parameter type | Example filter condition | | :--------------------------------- | :------------------------------------------------------------------------------------------------------------------------ | | String | `sensor_id = {{ var('sensor') }}` | | Date | `from_utc_timestamp(timestamp_col, 'UTC') > {{ var('start_date') }}` | | Numeric (Int, Long, Float, Double) | `total_usage_mb > {{ var('usage_cap_mb') }}` | | Array | Use `array_contains` in the visual expression builder. See [Use parameters](/data-analysis/development/parameters/usage). | | Boolean | `archived = {{ var('include_archived') }}` | For a full walkthrough of each parameter type including dashboard integration, see [Use parameters](/data-analysis/development/parameters/usage). ## Common issues ### No rows returned Verify that: * the filter condition matches the column data type * the values exist in the dataset * string comparisons use the expected capitalization ### Null values not matching Comparisons with `NULL` may return unexpected results. Use: * `IS NULL` * `IS NOT NULL` instead of: * `= NULL` * `!= NULL` ### Date comparisons not working Ensure date values use the correct format and data type. ## Similar tools and concepts The Filter gem works similarly to a SQL `WHERE` clause. For example, this filter condition: ```text theme={null} OrderAmount > 1000 ``` is equivalent to: ```sql theme={null} SELECT * FROM Orders WHERE OrderAmount > 1000 ``` If you've used SQL before, you can apply similar comparison operators and expressions in the Filter gem. You may also recognize similar behavior from: * the Alteryx Filter tool * `df.filter()` in PySpark * boolean indexing in Pandas ## Filter gem vs Conditional gem The Filter gem and [the Conditional gem](/data-analysis/gems/custom/condition) both evaluate conditions on a dataset, but they serve different purposes. ### Key differences | | Filter gem | Conditional gem | | :------------ | :------------------------------------------------------- | :-------------------------------------------------- | | Purpose | Reduce data | Route data | | Outputs | One output | Two or more outputs | | Behavior | Keeps rows that match the condition and removes the rest | Sends rows to different outputs based on conditions | | Routing | No routing | Yes | | Order matters | No | Yes (first matching output wins) | ### Row-level vs dataset-level behavior * The Filter gem always operates at the row level. * The Conditional gem can operate at: * row level (for example, `OrderAmount > 1000`) * dataset level (for example, `Count < threshold`) ### When to use the Filter gem Use the Filter gem when you want to: * keep only matching rows * remove unwanted records * reduce the size of a dataset * apply filtering logic similar to a SQL `WHERE` clause ### When to use the Conditional gem Use the Conditional gem when you want to: * route rows to different outputs * create branching logic * split data into multiple paths For example: * rows matching `OrderAmount > 1000` can be routed to `out0` * all remaining rows can be routed to `out1` ### Summary * Filter gem β†’ keeps matching rows * Conditional gem β†’ routes rows to different outputs Use the Filter gem for simple row filtering. Use the Conditional gem when you need branching or control flow in your pipeline. # FindDuplicates Source: https://docs.prophecy.ai/data-analysis/gems/prepare/find-duplicates Return rows that match a certain count This gem runs in . ## Overview The FindDuplicates gem filters rows in a dataset based on how frequently they appear. You can configure the gem to return the first occurrence of each unique group, identify groups that have duplicates, or apply custom filters based on group frequency or row position within groups. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Input and Output The FindDuplicates gem accepts the following input and generates one output. | Port | Description | | ------- | ------------------------------------------------------------------------------------------ | | **in0** | Input dataset to evaluate for duplicates. | | **out** | Output dataset with rows filtered according to the selected uniqueness and count criteria. | ## Parameters Configure the FindDuplicates gem using the following parameters. ### Unique rows generation scope Choose how to identify duplicates in your dataset. * **Uniqueness across all columns**: Considers the entire row when checking for uniqueness. Use this option when you want to find rows that are completely identical across every column. * **Uniqueness across selected columns**: Considers only specific columns when checking for uniqueness. Use this option when you want to find duplicates based on specific business keys or field combinations. ### Configure Grouping and Sorting Complete this step only if you selected **Uniqueness across selected columns**. * **Group By Columns**: Select the columns that determine whether rows are considered duplicates. For example, to find duplicate customer accounts, select the `email` column. * **Order rows within each group (Optional)**: Select columns to sort by and specify the sort order (ascending or descending). When multiple duplicates exist, this controls which row appears first in the results. ### Output records selection strategy Choose what to include in the output results. For each strategy, if you did not define a group, the dataset is grouped by all columns (groups consist of complete row matches). #### Unique Returns the first row in each group. All other rows in the group are considered duplicates and are excluded. #### Duplicate Returns all rows in a group except the first one. This is the inverse of Unique. #### Custom group count Filters rows based on how many times each group appears. Select a filter type: * **Group count equal to**: Returns rows that appear exactly `n` times * **Group count less than**: Returns rows that appear fewer than `n` times * **Group count greater than**: Returns rows that appear more than `n` times * **Group count not equal to**: Returns rows whose count differs from `n` * **Group count between**: Returns rows with count within a specified range In the **Grouped count** field, specify the count value `n`. #### Custom row number Filters rows based on position within each group. Select a filter type: * **Row number equal to**: Returns only the row with the specified position * **Row number less than**: Returns rows with a position less than `n` * **Row number greater than**: Returns rows with a position greater than `n` * **Row number not equal to**: Returns all rows except the one with the specified position * **Row number between**: Returns rows with a position within a specified range In the **Row number** field, specify the position value. Row numbering starts at 1 for each group. ## Example: Custom group count Given the following product catalog dataset:
| `product_id` | `category` | `region` | `last_updated` | | ------------ | ---------- | -------- | -------------- | | 001 | laptop | us-east | 2024-01-01 | | 002 | tablet | us-east | 2024-01-02 | | 003 | laptop | us-east | 2024-01-15 | | 004 | laptop | us-west | 2024-01-03 | | 005 | laptop | us-east | 2024-01-20 |
If you configure the gem with the following settings: * **Group by**: `category`, `region` * **Order by**: `last_updated` (descending) * **Output strategy**: Custom group count equal to 3 The gem creates the following groups: * `laptop` + `us-east`: 3 occurrences * `tablet` + `us-east`: 1 occurrence * `laptop` + `us-west`: 1 occurrence The output returns rows `005`, `003`, and `001` (laptop in us-east group) because this group appears exactly 3 times. The rows are ordered by most recent update date first. ## Example: Custom row number Given the following sales transactions dataset:
| `transaction_id` | `customer_id` | `date` | `amount` | | ---------------- | ------------- | ---------- | -------- | | T001 | C101 | 2024-01-05 | 120.00 | | T002 | C101 | 2024-02-10 | 250.00 | | T003 | C102 | 2024-01-07 | 75.00 | | T004 | C101 | 2024-03-15 | 300.00 | | T005 | C102 | 2024-02-20 | 150.00 |
If you configure the gem with the following settings: * **Group by**: `customer_id` * **Order by**: `date` (descending) * **Output strategy**: Row number greater than 1 The gem creates the following groups: * `C101`: 3 occurrences * `C102`: 2 occurrences The output keeps only rows whose position in the sorted order within their group is greater than 1.
| `transaction_id` | `customer_id` | `date` | `amount` | | ---------------- | ------------- | ---------- | -------- | | T002 | C101 | 2024-02-10 | 250.00 | | T001 | C101 | 2024-01-05 | 120.00 | | T003 | C102 | 2024-01-07 | 75.00 |
This output excludes the first (most recent) transaction from each customer group and returns all remaining transactions. Groups with only one row would not appear in the results. # FlattenSchema gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/prepare/flatten-schema Flatten nested columns This gem runs in . ## Overview Flattening a dataset schema helps you simplify complex, hierarchical data. Use the FlattenSchema gem to convert arrays or other nested data types into flat or lateral columns. This page describes how to use this gem according to your SQL warehouse provider. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Variant data When you import tables with variant data into Prophecy, you'll see one or more nested columns in your dataset schema. These will appear as a [variant data type](/data-analysis/gems/visual-expression-builder/variant-schema), which is an array of values with more than one data type. Variant types can be flattened. This data type is specific to Databricks and Snowflake. ## Parameters Configure the following parameters for the FlattenSchema gem. | Parameter | Description | Required | | ----------- | ---------------------------------------------------------------------- | -------- | | Flatten | Select the columns that include the arrays you want to flatten. | True | | Expressions | Define the output columns that you want to appear in the output table. | True | ## Example Assume you have the following JSON file that includes data you would like to flatten. The data includes information about multiple businesses with nested contact information for each business. Databricks accepts JSON files with a top-level array of JSON objects. ```json theme={null} [ { "first_name": "Remmington", "last_name": "Smith", "age": "68", "business": [ { "address": [ { "manager": "Liara Andrew", "name": "RS Enterprises", "contact": [ { "content": "rsmith@example.com", "type": "email" }, { "content": "1234-345-56", "type": "phone" } ], "is_still_active": false } ] } ] }, { "first_name": "Penny", "last_name": "John", "age": "57", "business": [ { "address": [ { "manager": "Bobby Frank", "name": "PJ Enterprises", "contact": [ { "content": "pjohn@example.com", "type": "email" }, { "content": "8203-512-49", "type": "phone" } ], "is_still_active": true } ] } ] } ] ``` Snowflake accepts JSON files with a top-level array of JSON objects. ```json theme={null} [ { "first_name": "Remmington", "last_name": "Smith", "age": "68", "business": [ { "address": [ { "manager": "Liara Andrew", "name": "RS Enterprises", "contact": [ { "content": "rsmith@example.com", "type": "email" }, { "content": "1234-345-56", "type": "phone" } ], "is_still_active": false } ] } ] }, { "first_name": "Penny", "last_name": "John", "age": "57", "business": [ { "address": [ { "manager": "Bobby Frank", "name": "PJ Enterprises", "contact": [ { "content": "pjohn@example.com", "type": "email" }, { "content": "8203-512-49", "type": "phone" } ], "is_still_active": true } ] } ] } ] ``` BigQuery expects newline-delimited JSON format. This means it expects one JSON object per line, without a surrounding array. ```json theme={null} {"first_name":"Remmington","last_name":"Smith","age":"68","business":[{"address":[{"manager":"Liara Andrew","name":"RS Enterprises","contact":[{"content":"rsmith@example.com","type":"email"},{"content":"1234-345-56","type":"phone"}],"is_still_active":false}]}]} {"first_name":"Penny","last_name":"John","age":"57","business":[{"address":[{"manager":"Bobby Frank","name":"PJ Enterprises","contact":[{"content":"pjohn@example.com","type":"email"},{"content":"8203-512-49","type":"phone"}],"is_still_active":true}]}]} ``` ### Expressions The FlattenSchema gem allows you to extract variant data into a flattened schema. For example, to flatten your variant data: 1. In the **Input** tab, hover over the `in0` field and click the **Add 12 Columns** button. Now, all the nested lowest-level values of your object are visible as columns in the `Expressions` section. Adding expressions 2. (Optional) To change the name of the column in the output, change the value in the `Output Column` for the row. 3. Click **Run**. ### Output After you run the FlattenSchema gem, click the **Data** button to see your schema based on the selected columns: Output interim The FlattenSchema gem flattened all your variant data, which gives you individual rows for each one. ## Snowflake advanced settings You can use advanced settings with your Snowflake source to customize the optional column arguments. To use the advanced settings: 1. Hover over the column you want to flatten. 2. Click the dropdown arrow. You can customize the following options: | Option | Description | Default | | ----------------------------------- | ------------------------------------------------------------------------------------------- | ------- | | Path to the element | Path to the element within the variant data structure that you want to flatten. | None | | Flatten all elements recursively | Whether to expand all sub-elements recursively. | `false` | | Preserve rows with missing field | Whether to include rows with missing fields as `null` in the key, index, and value columns. | `false` | | Datatype that needs to be flattened | Data type that you want to flatten. Possible values are: `Object`, `Array`, or `Both`. | `Both` | # GenerateRows Source: https://docs.prophecy.ai/data-analysis/gems/prepare/generate-rows Create new rows of data using iterative expressions This gem runs in . ## Overview The GenerateRows gem creates new rows of data at the record level using iterative expressions. This gem generates sequences of numbers, dates, or other values based on initialization, condition, and loop expressions. The gem follows a three-step process to create rows: 1. **Initialization Expression**: Sets the starting value for the first row. 2. **Condition Expression**: Defines when to stop generating rows (true/false condition). 3. **Loop Expression**: Specifies how values change between iterations. The gem continues generating rows until the condition expression evaluates to false, at which point the generation process terminates. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Input and Output The GenerateRows gem accepts optional input data and produces one output dataset containing the generated rows. | Port | Description | | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **in0** | Optional input dataset. Columns from this dataset can be referenced in the gem expressions.
  • If an input dataset is provided, the gem generates a separate sequence of rows for each input row.
  • If no input is provided, the gem generates a single sequence.
| | **out** | Output dataset containing:
  • All original input columns (if input is provided)
  • The new generated column with values created by the iterative expressions
Each input row may produce multiple output rows depending on the loop and condition expressions. | ## Parameters Configure the GenerateRows gem using the following parameters. | Parameter | Description | | --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Choose output column name | Name of the column that will contain the generated values. Each generated value in this column corresponds to one row created by the gem’s iterative process. | | Configure row generation strategy | SQL logic for generating rows using three expressions:
  • **Initialization expression**: Sets the starting value for the generation process. This is the value of the first generated row.
  • **Condition expression**: Determines how long to continue generating rows. Rows are added while this expression evaluates to true. Can reference input columns or the output column value.
  • **Loop expression (usually incremental)**: Specifies how the generated value changes each iteration. Typically an increment, but can be any expression that modifies the output column value.
| | Max rows per iteration | Maximum number of rows generated per sequence. This limit prevents infinite loops or excessive memory usage.
**Default**: 100,000 rows. | To reference input columns in the **Condition expression** and **Loop expression** fields, you must use the following format: ``` payload. ``` ### Expression examples Reference the following examples to understand how to build different expressions for this gem. All expressions must be valid SQL. To reference the iterative value, use the output column name you defined. #### Initialization expression examples | Expression | Description | | ---------------- | -------------------------------------------------- | | `1` | Start with the number 1 | | `StartDate` | Use a value from an input column named `StartDate` | | `CURRENT_DATE()` | Start with today's date | #### Condition expression examples | Expression | Description | | ------------------ | ------------------------------------------------------------------------- | | `value <= 10` | Generate rows while value is less than or equal to 10 | | `value < MaxValue` | Continue until value reaches a maximum from input column named `MaxValue` | | `date <= EndDate` | Generate dates until reaching an end date defined in an `EndDate` column | #### Loop expression examples | Expression | Description | | ----------------------- | ------------------------------------------------------------------------ | | `value + 1` | Increment by 1 each iteration | | `value + StepSize` | Increment by a variable step size defined in the input column `StepSize` | | `date + INTERVAL 1 DAY` | Add one day for date sequences | ## Gem examples The following sections demonstrate different gem configurations and their corresponding outputs. ### Create sequence Assume you need to create a sequence of numbers from 1 to 10 for testing purposes or to generate row IDs. You want to create a simple numeric sequence that can be used as a foundation for other data operations. Use the following configuration: * **Input dataset**: none * **Choose output column name**: `value` * **Initialization expression**: `1` * **Condition expression**: `value <= 10` * **Loop expression (usually incremental)**: `value + 1` * **Max rows per iteration**: `100` If no input table is present, the following output table is generated:
| value | | ----- | | 1 | | 2 | | 3 | | 4 | | 5 | | 6 | | 7 | | 8 | | 9 | | 10 |
The gem generates 10 rows with sequential values from 1 to 10. The process starts with the initialization value of 1, continues generating rows while the condition `value <= 10` is true, and increments the value by 1 in each iteration. ### Increment dates Assume you want to generate a range of dates. To do so, you can use the following gem configuration: * **Input dataset**: none * **Choose output column name**: `date` * **Initialization expression**: `'2024-01-01'` * **Condition expression**: `date <= '2024-01-07'` * **Loop expression (usually incremental)**: `date + INTERVAL 1 DAY` This would generate 7 rows with consecutive dates from January 1st through January 7th, 2024. ### Leverage input data Assume you have input data that you want to use to influence the expressions in the gem configuration. You might have the following input table called `Constants`:
| `min_value` | `max_value` | `step_size` | | ----------- | ----------- | ----------- | | 5 | 20 | 3 |
To leverage this table, you can use the following gem configuration: * **Input dataset**: `Constants` table * **Choose output column name**: `sequence` * **Initialization expression**: `min_value` * **Condition expression**: `sequence <= max_value` * **Loop expression (usually incremental)**: `sequence + step_size` This would produce the following output table:
| `min_value` | `max_value` | `step_size` | `sequence` | | ----------- | ----------- | ----------- | ---------- | | 5 | 20 | 3 | 5 | | 5 | 20 | 3 | 8 | | 5 | 20 | 3 | 11 | | 5 | 20 | 3 | 14 | | 5 | 20 | 3 | 17 | | 5 | 20 | 3 | 20 |
The gem processes the input row and generates a sequence starting at 5, incrementing by 3, and continuing until reaching 20. # Imputation Source: https://docs.prophecy.ai/data-analysis/gems/prepare/imputation Replace specified values in numeric fields with calculated or user-defined replacement values This gem runs in . You can use the Imputation gem to replace a specified value in one or more numeric fields with a replacement value before downstream analysis or modeling. For example, you can replace null values in `sales `fields with the average of non-null `sales` so those missing values do not distort later calculations. ## Prerequisites Add `prophecy_basics` package version 1.0.11 or higher to your project. ## Parameters | Parameter | Description | | ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Fields to impute | Select one or more numeric fields to update. | | Incoming value to replace | Choose which value should be replaced in the selected fields.
  • Null
  • User-specified value
If you choose User-specified value, the **Value to Replace** field appears. | | Value to Replace | Enter the value to replace when **Incoming value to replace** is set to User-specified value. | | Replace with value | Choose the value used for replacement.
  • Average
  • Median
  • Mode
  • User-specified value
If you choose User-specified value, the **Replacement value** field appears. | | Replacement value | Enter the replacement value when **Replace with value** is set to User-specified value. | | Include imputed value indicator field | Add an indicator field for each imputed column that shows whether a value was imputed. | | Output imputed values as a separate field | Keep the original field unchanged and write the imputed result to a new column. (Non-imputed values are automatically included in column.) | ## How it works The Imputation gem scans the selected fields and looks for values that match the configured **Incoming value to replace** setting. Matching values are replaced using the selected method: * **Average**: Replaces with the mean of valid values in the field, excluding the value being replaced. * **Median**: Replaces with the middle value in the field, excluding the value being replaced. * **Mode**: Replaces with the most frequently occurring value in the field, excluding the value being replaced. * **User-specified value**: Replaces with the value you provide. ## Output By default, the output contains the original data stream with imputed values written back into the selected fields. When **Include imputed value indicator field** is enabled, an additional field is added for each imputed field to indicate whether the value was imputed. Naming pattern: `_Indicator` When **Output imputed values as a separate field** is enabled, the original field is preserved and a new field is added with the imputed result. Naming pattern: `_ImputedValue` If both options are selected, both additional fields are included. ## Notes * This gem works for numeric fields. * Imputation is calculated separately for each selected field. * When using Average, Median, or Mode, the replacement statistic is calculated using valid values only, excluding the value being replaced. * If you do not select **Output imputed values as a separate field**, the original field is overwritten with the imputed result. ## Example Suppose you have the following dataset with missing values in numeric fields.
| Product | Price | Total\_Sale | | ------- | ----- | ----------- | | Shirt | 20.0 | 200.0 | | Pants | null | 150.0 | | Jacket | 50.0 | null | | Shoes | 30.0 | 300.0 |
If you configure the gem to replace **Null** values using **Average**, the null for "Jacket" is replaced with the average of `Total_Sale`. ### Result
| Product | Price | Total\_Sale | | ------- | ----- | ----------- | | Shirt | 20.0 | 200.0 | | Pants | 33.3 | 150.0 | | Jacket | 50.0 | 216.7 | | Shoes | 30.0 | 300.0 |
# Limit gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/prepare/limit Limit the number of columns processed This gem runs in . ## Overview Limits the number of rows in the output. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ### Parameters | Parameter | Description | Required | | :-------- | :---------------------------------------------------------------------------------- | :------- | | Input | The dataset to limit. | True | | Limit | Number of rows required in output.
**Allowed range**: \[0, 231 -1] | True | # DataMasking Source: https://docs.prophecy.ai/data-analysis/gems/prepare/masking Obfuscate data in one or more columns This gem runs in . ## Overview The DataMasking gem allows you to obfuscate sensitive string data in one or more columns. This can be useful for protecting personally identifiable information (PII) or other confidential values. This page describes the available masking methods in the gem and other configuration options. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Input and Output The DataMasking gem accepts the following input and output. | Port | Description | | ------- | -------------------------------------------------------------------------------------------------------------------------------- | | **in0** | Input dataset containing columns with data you want to mask. | | **out** | Output dataset with the masked data. The output schema depends on the [masked column option](#masked-column-options) you choose. | ## Parameters Configure the DataMasking gem using the following parameters. | Parameter | Description | | ---------------------------------- | ------------------------------------------------------------------------------------------------------------ | | Select columns to apply masking on | One or more string-type columns from the input dataset to transform. Only `string` columns are supported. | | Masking method | The method used to obfuscate the data. Jump to [Masking methods](#masking-methods) to learn more. | | Masked column options | Choose how to output the masked data. Jump to [Masked column options](#masked-column-options) to learn more. | ### Masked column options #### Substitute the new columns in place The new values will appear in the original column. #### Add new columns with a prefix/suffix attached The new values will appear in a new column that will have the original name with a prefix or suffix that you specify. #### Apply a single hash to all the selected columns at once There will only be one new output column, regardless of the number of columns to apply masking on. This only applies to the `hash` masking method. ### Masking methods Choose one of the following techniques to obfuscate string data. Some methods support additional configuration options. #### `mask` Replaces characters in each string with substitute characters based on character type. Applies individually to each selected column. This method lets you optionally define the following additional parameters: | Name | Description | | ------------------------- | ---------------------------------------------------------------------------------------------------------------- | | Upper char substitute key | Character to replace uppercase letters. Default is `'X'`. Use `NULL` to keep the original. | | Lower char substitute key | Character to replace lowercase letters. Default is `'x'`. Use `NULL` to keep the original. | | Digit char substitute key | Character to replace digits. Default is `'n'`. Use `NULL` to keep the original. | | Other char substitute key | Character to replace all other characters. The default is `NULL`, which leaves the original characters unmasked. | #### `hash` Applies a hash function to column values. This method will apply the hash in one of two ways: * The hash function is applied to each input column individually, and a new output column will be created for each input column. * The hash function is applied to all selected columns combined into a single hash value in a single new output column. #### `crc32` Applies the CRC32 hash function to each selected column. This method has no additional parameters. #### `sha` Applies the SHA-1 hash function to each selected column. This method has no additional parameters. #### `sha2` Applies the SHA-2 hashing algorithm to each selected column. This method lets you select the bit length for masking: * Bit length can be `224`, `256`, `384`, or `512`. #### `md5` Applies the MD5 hash function to each selected column. This method has no additional parameters. ## Example Assume you have a table that you would like to mask using the `mask` method. Using the default mask parameters, the string `John.Doe123!` will be converted to `Xxxx.Xxxnnn!`. # MultiColumnEdit Source: https://docs.prophecy.ai/data-analysis/gems/prepare/multi-column-edit Change the data type of multiple columns at once This gem runs in . ## Overview The MultiColumnEdit gem primarily lets you cast or change the data type of multiple columns at once. It provides additional functionality, including: * Adding a prefix or suffix to selected columns. * Applying a custom expression to selected columns. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Parameters | Parameter | Description | Required | | --------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | | Selected columns to edit | Choose which columns you want to transform. | Yes | | Maintain the original columns and add prefix/suffix to the new column | If checked, the original columns will stay the same and new ones will be added with the prefix/suffix. Otherwise, the original column names change. | No | | Prefix/Suffix dropdown | Lets you choose whether to add text at the beginning (prefix) or end (suffix) of the column names. | No | | Build a single expression to apply to all selected columns | A SQL expression you apply to each selected column. If you don't want to change the values, enter `column_value`. | Yes | # MultiColumnRename Source: https://docs.prophecy.ai/data-analysis/gems/prepare/multi-column-rename Quickly standardize column names using a consistent naming pattern This gem runs in . ## Overview Use the MultiColumnRename gem to efficiently rename several columns in your dataset at the same time. This gem allows you to apply a consistent naming pattern by adding a prefix or suffix to selected columns, or perform advanced renaming using custom SQL expressions. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Parameters | Field | Description | | ------------------------ | ------------------------------------------------------------------------------------------------------------------ | | Select columns to rename | Set of columns that you will rename. | | Rename method | How you will rename columns.
You can either add a prefix/suffix, or choose advanced rename (SQL expression). | ## Example Assume you have the following table that includes the weather forecast for the next four days.
| DatePrediction | TemperatureCelsius | HumidityPercent | WindSpeed | Condition | | -------------- | ------------------ | --------------- | --------- | --------- | | 2025-03-01 | 15 | 65 | 10 | Sunny | | 2025-03-02 | 17 | 70 | 12 | Cloudy | | 2025-03-03 | 16 | 68 | 11 | Rainy | | 2025-03-04 | 14 | 72 | 9 | Sunny |
To standardize column names by converting them to lowercase, use the **Advanced rename** option in the MultiColumnRename gem with a custom SQL expression. 1. Create a **MultiColumnRename** gem. 2. Open the gem configuration and stay in the **Visual** view. 3. Under **Select columns to rename**, select all columns. 4. For the **Rename method**, choose **Advanced rename**. 5. Click **Select expression > Function**. 6. Search for and select the `lower` function. 7. Inside of the `lower` function, click **expr > Custom Code**. 8. Inside of the code box, write `column_name`. This applies the function to the column name. 9. Click **Done** on the code box, and then click **Save** on your gem. ### Result After saving and running the gem, all selected columns will be renamed using the lower function. In this case, all column names will be lowercase in the output table.
| `dateprediction` | `temperaturecelsius` | `humiditypercent` | `windspeed` | `condition` | | ---------------- | -------------------- | ----------------- | ----------- | ----------- | | 2025-03-01 | 15 | 65 | 10 | Sunny | | 2025-03-02 | 17 | 70 | 12 | Cloudy | | 2025-03-03 | 16 | 68 | 11 | Rainy | | 2025-03-04 | 14 | 72 | 9 | Sunny |
# OrderBy gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/prepare/order-by Sort the data This gem runs in . ## Overview Sorts a model on one or more columns in ascending or descending order. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Parameters | Parameter | Description | Required | | ------------- | ------------------------------------------ | -------- | | Order columns | Columns to sort the model by | True | | Sort | Order of sorting (ascending or descending) | True | ## Example Assume you have the following weather prediction table.
| DatePrediction | TemperatureCelsius | HumidityPercent | WindSpeed | Condition | | -------------- | ------------------ | --------------- | --------- | --------- | | 2025-03-01 | 15 | 65 | 10 | Sunny | | 2025-03-02 | 17 | 70 | 12 | Cloudy | | 2025-03-03 | 16 | 68 | 11 | Rainy | | 2025-03-04 | 14 | 72 | 9 | Sunny |
### Result The follow table results when you order by the `HumidityPercent` column in **ascending** order.
| DatePrediction | TemperatureCelsius | HumidityPercent | WindSpeed | Condition | | -------------- | ------------------ | --------------- | --------- | --------- | | 2025-03-01 | 15 | 65 | 10 | Sunny | | 2025-03-03 | 16 | 68 | 11 | Rainy | | 2025-03-02 | 17 | 70 | 12 | Cloudy | | 2025-03-04 | 14 | 72 | 9 | Sunny |
# RecordID Source: https://docs.prophecy.ai/data-analysis/gems/prepare/record-id Assign each row of a table a unique ID This gem runs in . Assigning unique identifiers to each row in a dataset is a common requirement for data preparation. The RecordID gem allows you to easily generate row-level IDs using two methods: * UUID for randomly generated values * Incremental ID for ordered, sequential values You can customize how IDs are added, including naming the column, setting the data type and format, and specifying where the column appears in the output schema. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Input and Output The RecordID gem accepts the following input and generates one output. | Port | Description | | ------- | --------------------------------------------------------------- | | **in0** | Input dataset containing the records you wish to assign IDs to. | | **out** | Output dataset with a new record ID column. | ## Parameters Review the following gem parameters by method. ### UUID The UUID method assigns a universally unique identifier (UUID) to each row. These values are randomly generated and are ideal when you need non-sequential, non-predictable IDs. | Parameter | Description | | ------------------ | ---------------------------------------------------------------------------------- | | Output Column Name | Name of the new column where the generated UUIDs will be stored. | | Column position | Choose to **add as first column** or **add as last column** in the output dataset. | ### Incremental ID The Incremental ID method generates sequential values starting from a specified number and increasing by 1 for each row. You can also group and sort the data to restart numbering within each group and control the order in which IDs are assigned. | Parameter | Description | | --------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Output Column Name | Name of the new column that will contain the generated record IDs. | | Starting Value | The first number in the sequence. | | Data type | Select **integer** or **string** as the output type of the record ID values. | | Size (String only) | Total number of characters in the string.
Leading zeros will be added if the starting value is shorter than the defined size. | | Column position | Choose to **add as first column** or **add as last column** in the output dataset. | | Record ID Generation Scope | Specify whether to generate IDs **across entire table** or **within each group** defined by selected columns. | | Group By Columns | When generating IDs within groups, choose one or more columns to group the data by. | | Order Rows Within Each Group (Optional) | (Optional) Define the columns to determine the order in which rows are numbered within each group. You can select multiple columns and sort them in ascending or descending order. | # Reformat gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/prepare/reformat Use expressions to reformat column names and values This gem runs in . ## Overview Use the Reformat gem to: * rename columns * create calculated columns * modify existing column values * change data types * select which columns appear in the output dataset Use expressions and functions to transform values, combine columns, clean data, or create new derived fields. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Common transformation examples | Goal | Example expression | | | | | | :-------------------------- | :----------------------------------------- | - | --- | - | ------------ | | Rename a column | Set `Target column` to the new column name | | | | | | Create a full name column | \`first\_name | | ' ' | | last\_name\` | | Convert text to uppercase | `UPPER(customer_name)` | | | | | | Replace null values | `COALESCE(region, 'Unknown')` | | | | | | Create a calculated field | `price * quantity` | | | | | | Convert a value to a string | `CAST(order_id AS STRING)` | | | | | | Extract part of a date | `YEAR(order_date)` | | | | | ## Parameters | Parameter | Description | Required | | :------------ | :---------------------------------- | :--------------------------------------- | | Model | Input dataset to transform | True | | Target column | Output column name | False | | Expression | Expression to compute target column | Required if a `Target column` is present | If no columns are selected, then all columns are passed through to the output. ### Preview transformed data Run the Reformat gem to generate preview results. After the initial run, the **Data** tab displays how your transformations affect the output dataset. As you edit expressions, the preview updates automatically so you can immediately verify the results of your changes. ### Visualize data flow You can preview the changes that the Reformat gem will make by either: 1. Clicking the the **Visualize Data Flow** icon to the right of an expression, or 2. Clicking the **Data** button at the bottom of the Reformat gem visual editor. When enabled: * Input columns referenced by the expression are highlighted in the input data preview. * The output column produced by the expression is highlighted in the output preview highlights update automatically as you edit the expression. This view helps you understand column dependencies and verify how output values are derived from source data. You must run the Reformat gem at least once before preview data and data flow visualizations are available. visualize data flow ## Common issues ### Column not found Verify that: * The column name exists in the input dataset. * The column name uses the correct capitalization. * The column reference is spelled correctly. ### Type mismatch errors Some functions and operators require specific data types. For example: * Numeric calculations require numeric columns. * String functions require text values. Use `CAST()` to convert values when needed. ### Null values causing unexpected results Some expressions return `NULL` when one or more input values are `NULL`. Use `COALESCE()` to replace null values with defaults. ### Duplicate column names Output column names must be unique. ## Similar tools and concepts The Reformat gem can be used to: * rename columns * create calculated fields * modify existing column values * select output columns You may recognize similar behavior from: * SQL `SELECT` expressions and aliases * the Alteryx Formula and Select tools * PySpark `select()` and `withColumn()` * Pandas column transformations # Tile Source: https://docs.prophecy.ai/data-analysis/gems/prepare/tile Assign tile values to records by using one of five tiling methods. This gem runs in . Use the Tile gem to assign records to tiles based on your data and the tiling method you choose. You can distribute records by equal sum, equal record count, standard deviation bands, unique values, or manually defined cutoffs. The gem appends tile information to your data so you can analyze ranges, sequence within each tile, and grouped distributions. Depending on the method, you can also group records first or control how rows are ordered before tiles are assigned. ## Prerequisites Add `prophecy_basics` package version 1.0.11 or higher to your project. ## Overview The Tile gem supports these tiling methods: * **Equal Sum** * **Equal Records** * **Smart Tile** * **Unique Value** * **Manual** The output appends two fields to the data: * **Tile number**, which stores the tile assigned to the record * **Tile sequence number**, which stores the record's position within the tile ## Parameters | Parameter | Description | | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | | Select tiling method | Choose the tiling method. Available options are **Equal Sum**, **Equal Records**, **Smart Tile**, **Unique Value**, and **Manual**. | ## Equal Sum Use **Equal Sum** to assign tiles so each tile has approximately the same total from the selected sum field based on the row order. ### Parameters | Parameter | Description | | --------------------------------------- | --------------------------------------------------------------------------------------------------------------- | | Number of Tiles | Specify how many tiles to assign. | | Select Sum Column | Select the numeric field used to distribute totals across tiles. | | Select Column | Select the column used for tiling configuration. | | Select group by columns (Optional) | Optionally select fields to create tiles independently within each group. | | Order rows within each group (Optional) | Optionally order rows before assigning tiles. This section includes **Order By Columns** and **Sort strategy**. | ### Order rows within each group Use **Order By Columns** to select the columns used to order rows. Use **Sort strategy** to choose one of the following options: * ascending nulls first * ascending nulls last * descending nulls first * descending nulls last ### Example Let's say you have the following dataset and set **Number of Tiles** to `2`.
| Record | Sales | | ------ | ----- | | A | 10 | | B | 20 | | C | 15 | | D | 15 |
If you select **Equal Sum** and use **Sales** as the sum column, the gem assigns records so each tile has a similar total based on the row order. ### Result
| Record | Sales | Tile number | Tile sequence number | | ------ | ----- | ----------- | -------------------- | | A | 10 | 1 | 1 | | B | 20 | 1 | 2 | | C | 15 | 2 | 1 | | D | 15 | 2 | 2 |
In this example, both tiles total `30`. ## Equal Records Use **Equal Records** to divide input records into the specified number of tiles so each tile contains the same number of records as closely as possible. ### Parameters | Parameter | Description | | --------------------------------------- | ------------------------------------------------------------------------- | | Number of Tiles | Specify how many tiles to assign. | | Select group by columns (Optional) | Optionally select fields to create tiles independently within each group. | | Do not split tile on Columns (Optional) | Optionally prevent a tile from splitting across the selected column. | | Order rows within each group (Optional) | Optionally order rows before assigning tiles. | ### Example Let's say you have the following dataset and set **Number of Tiles** to `2`.
| Record | Score | | ------ | ----- | | A | 90 | | B | 85 | | C | 80 | | D | 75 |
If you select **Equal Records**, the gem divides the four records into two tiles with two records in each tile. ### Result
| Record | Score | Tile number | Tile sequence number | | ------ | ----- | ----------- | -------------------- | | A | 90 | 1 | 1 | | B | 85 | 1 | 2 | | C | 80 | 2 | 1 | | D | 75 | 2 | 2 |
## Smart Tile Use **Smart Tile** to create tiles based on the standard deviation of values in a numeric field. The assigned tile indicates whether the value falls within the average range, above the average, or below the average. ### Parameters | Parameter | Description | | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | | Select Tile Numeric Column | Select the numeric field used for the tile calculation. | | Select Column | Select the column used for tiling configuration. | | Select group by columns (Optional) | Optionally select fields to create tiles independently within each group. | | Do not output name column | Do not append a descriptive name column. | | Output name column | Append a descriptive output field. Descriptors include Average, Above Average, High, Extremely High, Below Average, Low, and Extremely Low. | | Output verbose name column | Append a descriptive output field and include the value range in parentheses. | ### Example Let's say you use **Price** as the tile numeric column and enable **Output name column**.
| Sale\_ID | Price | Tile number | SmartTile\_Num | Tile sequence number | | -------- | ----- | ----------- | -------------- | -------------------- | | 2 | 11.89 | -1 | Below Average | 1 | | 12 | 15.16 | -1 | Below Average | 2 | | 4 | 18.06 | -1 | Below Average | 3 | | 27 | 19.91 | -1 | Below Average | 4 |
In this example, the records fall into the `-1` band and the output includes the descriptive label **Below Average**. ## Unique Value Use **Unique Value** to assign a unique tile to each unique value in one or more selected fields. If you select multiple fields, the tile is based on the combination of values. ### Parameters | Parameter | Description | | ---------------------------------- | ------------------------------------------------------------------------- | | Select Unique Column | Select one or more fields used to assign unique tiles. | | Select group by columns (Optional) | Optionally select fields to create tiles independently within each group. | ### Example Let's say you have the following dataset and select **Category** as the unique column.
| Record | Category | | ------ | -------- | | A | Books | | B | Games | | C | Books | | D | Music |
If you select **Unique Value**, each unique category receives its own tile. ### Result
| Record | Category | Tile number | Tile sequence number | | ------ | -------- | ----------- | -------------------- | | A | Books | 1 | 1 | | C | Books | 1 | 2 | | B | Games | 2 | 1 | | D | Music | 3 | 1 |
## Manual Use **Manual** to define tile cutoffs yourself. ### Parameters | Parameter | Description | | ------------------------------ | -------------------------------------------------- | | Select Tile Numeric Column | Select the numeric field used for tile assignment. | | Enter one or more tile cutoffs | Enter each tile's upper limit separated by commas. | ### Example Let's say you have the following dataset, select **Score** as the tile numeric column, and enter the cutoffs `50,80`.
| Record | Score | | ------ | ----- | | A | 30 | | B | 60 | | C | 90 |
If you select **Manual**, the gem assigns tiles based on the cutoffs you provide. ### Result
| Record | Score | Tile number | Tile sequence number | | ------ | ----- | ----------- | -------------------- | | A | 30 | 1 | 1 | | B | 60 | 2 | 1 | | C | 90 | 3 | 1 |
# Email report gem Source: https://docs.prophecy.ai/data-analysis/gems/report/email Send your output tables from your pipeline to others via email This gem runs in . ## Overview Use the Email gem to send output tables from your pipeline to others via email. You can configure static values for recipients, subject, and body, or dynamically populate these fields from columns in your input dataset using the **Use column** option. ## Input The Email gem supports one or two input ports, depending on how you want to provide email content and attachments. * **Single input port**: The dataset can be used for both email content (To, Cc, Bc, Subject, Body) and attachment data (if enabled). * **Two input ports**: The first input port provides email content and parameters, while the second port provides attachment data. The gem will send email(s) when it runs. No output table will be written to your data warehouse. ## Parameters Review the following Email gem parameters. Parameters can be set manually or configured dynamically using the **Use column** option. When you enable the **Use column** checkbox, the gem field is automatically populated with the value from the specified column in your input dataset. | Parameter | Description | Use column checkbox | | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------- | | Select or create connection | Defines which [SMTP connection](/data-analysis/environment/connections/smtp) to use for the gem.
This will also determine the **sender** of the email. | Not applicable | | To | Specifies the recipient(s) of the email. You can add multiple recipients. | Supported | | Cc | Specifies the recipients to be included in the CC field. | Supported | | Bc | Specifies the recipients to be included in the BCC field. | Supported | | Subject | Defines the subject of the email, providing a brief summary of its content. | Supported | | Body | Contains the main content or body of the email, where you can provide the message.
Enable the **Use Custom HTML** to paste HTML directly into the email body. | Supported | | Include Data as Attachment | Checkbox that enables sending input data as an attachment in the email.
You can send the data as an XLSX or CSV file.
File extensions are not required when naming the file. | Not applicable | ## Example Suppose you have a dataset that you want to sent to different recipients. * **Input port 1**: Email parameters dataset
| Recipients | Body | | ------------------- | --------------------------------------------- | | `team1@example.com` | Hello Team 1, please see the attached report. | | `team2@example.com` | Hello Team 2, please see the attached report. | | `team3@example.com` | Hello Team 3, please see the attached report. |
* **Input port 2**: Attachment dataset To send the attachment data to the recipients in the table: 1. Connect the email parameters dataset to port 1 of the Email gem. 2. Connect the attachment dataset to port 2 of the Email gem. 3. Open the Email gem. 4. Enable the **Use column** checkbox for **To** and **Body** fields. 5. Select the columns from input 1 that correspond to the email parameters. 6. Enable **Include Data as Attachment** to attach the dataset as an XLSX or CSV file. # PowerBIWrite Source: https://docs.prophecy.ai/data-analysis/gems/report/power-bi Send your pipeline output directly to PowerBI This gem runs in . ## Overview The PowerBIWrite gem lets you publish pipeline results directly to Power BI tables. This gem supports fine-grained options like write modes and schema management to control how tables are written. You can configure the gem to either write tables to new datasets or existing ones in a specified Power BI workspace. ## Inputs The PowerBIWrite gem accepts the following inputs. | Port | Description | | ------- | ------------------------------------------------- | | **in0** | Table to add or update in the dataset. | | **inN** | Additional table to add or update in the dataset. | To add additional input ports, click `+` next to **Ports**. ## Parameters Use the following parameters to configure the PowerBIWrite gem. | Parameter | Description | | ---------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Power BI [connection](/data-analysis/environment/connections/power-bi) to use for the gem. | | Workspace Name | Power BI [workspace](https://learn.microsoft.com/en-us/power-bi/collaborate-share/service-new-workspaces) that contains or will contain the dataset. | | Create New or Use Existing Dataset | Choose **Dataset Name** to create a new dataset in the workspace.
Choose **Dataset ID** to push tables to an existing dataset in the workspace.
These options are described in detail in the following sections. | **Datasets** in the PowerBIWrite gem refer to the *semantic model* content type in Power BI. For more information, visit [New name for Power BI datasets](https://learn.microsoft.com/en-us/power-bi/connect-data/service-datasets-rename). ### Dataset Name Select this option to create a new dataset in your workspace. You will need to give the dataset a name that will appear in Power BI. #### Table Write Configuration This configuration lets you define how your table(s) will be written to the dataset. Each row accepts the following parameters: | Parameter | Description | | ----------- | ---------------------------------------------------------------------- | | Input Alias | Input port that maps to a table in Power BI. Example: **in0**, **in1** | | Table Name | Write a corresponding table name that will appear in Power BI. | ### Dataset ID Select this option to update table inside an existing dataset in your workspace. You will need the Dataset ID to identify the existing dataset. The Dataset ID is typically part of the URL when you open a dataset in Power BI. #### Table Write Configuration This configuration lets you define how your table(s) will be written to the dataset. Each row accepts the following parameters: | Parameter | Description | | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Input Alias | Input port that maps to a table in Power BI. Example: **in0**, **in1** | | Table Name | Write a corresponding table name that will appear in Power BI. | | Write Mode | Controls how data is written to the Power BI table.
Choose **Append** when you want to preserve existing data and continuously add new entries.
Choose **Overwrite** if you want to fully refresh the table in Power BI. | | Overwrite Schema | Determines whether the schema (columns and their types) in Power BI should be replaced when it differs from the incoming dataset.
Choose **Yes** if you expect the schema to evolve and want Power BI to reflect those changes automatically.
Choose **No** to preserve the current schema in Power BI, even if the incoming data has a different structure. | # TableauWrite Source: https://docs.prophecy.ai/data-analysis/gems/report/tableau Send data to automatically update your Tableau dashboards This gem runs in . ## Overview The TableauWrite gem lets you send data that updates the data sources that are utilized by Tableau dashboards. ## Input The TableauWrite gem accepts the following inputs. | Port | Description | | ------- | ------------------------------------------------------------------------------------------------------------------- | | **in0** | The table that will be sent as a `Hyper` file to update your Tableau data source. You can only configure one input. | No output table will be written to your data warehouse. ## Parameters | Parameter | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Defines which [Tableau connection](/data-analysis/environment/connections/tableau) to use for the gem. | | Project name | The name of your Tableau project that contains the data source to update. | | Data source | The data source you want to update in Tableau.
A data source is a connection to data that you are using for analysis and visualization. | Tableau Gem configuration # Visualize Source: https://docs.prophecy.ai/data-analysis/gems/report/visualize Add checkpoints to your pipeline to capture data for analysis dashboards The Visualize gem doesn't perform any data transformations. Instead, it marks a point in your pipeline where you want to capture data to display in an [analysis dashboard](/data-analysis/analysis/overview). This gem lets you report data outcomes without having to materialize the data in your data warehouse or another data storage location. ## Input and output The Visualize gem passes through input data exactly as is. ## Parameters Open the Visualize gem configuration and select the analysis to add the data to. There are other ways to accomplish this, depending on your workflow: * When building the analysis, you can choose any existing Visualize gem from the pipeline for data integration components. * As you chat with the Prophecy Agent, you can prompt it to add the data from the Visualize gem to an analysis. In all cases, the Visualize gem **must be present** in the pipeline. ## Data persistence Because the Visualize gem does not persist or materialize data, you must run the analysis (running the pipeline) to populate data components in the dashboard. **The data isn't saved anywhere.** To persist data without running the pipeline, use a [Table gem](/data-analysis/gems/source-target/source-target) instead of a Visualize gem. This way, Prophecy writes the data to your data warehouse and can use that table as the source for the dashboard. Evaluate your use case to determine whether it's more beneficial to materialize the data (higher storage cost) or run the pipeline each time (higher compute cost). # Add data sources Source: https://docs.prophecy.ai/data-analysis/gems/source-target/adding-data-sources Use the Environment tab to quickly add tables and files to your pipeline The easiest way to add data sources to your pipeline is through the **Environment** tab in the left sidebar. This method automatically configures the gem settings, so you can start using the data immediately without manually creating and configuring Source and Table gems. ## Add data from the Environment tab The Environment tab displays all tables and files available through [connections](/data-analysis/environment/connections/connections) defined in your attached fabric. You can browse by connection type or use the search bar to find specific datasets. 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click the **Environment** tab in the left sidebar. 3. Expand the connection that contains your data source. You can add data to your pipeline using either method: * **Drag and drop**: Drag the table or file onto the canvas. * **Add button**: Hover over the table or file and click **Add**. The gem appears on your canvas with all configuration settings already applied. The gem type depends on where your data is stored: * **Table gems**: For datasets from your [SQL warehouse](/data-analysis/environment/fabrics/prophecy-fabrics) * **Source gems**: For tables or files from external systems ## Modify gem configuration After adding a data source from the Environment tab, you can adjust any configuration settings: 1. Click the gem on the canvas to open its configuration panel. 2. Modify location paths, schema definitions, format properties, or connection settings as needed. 3. Save your changes. Changes you make to the gem configuration only affect that specific gem instance. The original data source remains unchanged. ## Alternative: Create gems from scratch If you need to create a gem manually, you can add Source, Target, or Table gems directly from the canvas: 1. Click **Source/Target** in the gem drawer at the top of the canvas. 2. Select the gem type from the dropdown. 3. Click the gem to open the configuration panel. 4. Configure all settings manually, including selecting connections, defining paths, and setting up schemas. Creating gems from scratch requires you to configure all settings manually. This approach is more time-consuming and error-prone than using the Environment tab, which automatically applies the correct configuration based on your existing connections. # Google BigQuery external table gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/bigquery Read and write catalog tables in BigQuery This gem runs in . ## Overview This page describes how to use BigQuery external Source and Target gems to read from or write to tables. Only use an external Source and Target gem when BigQuery is not the configured [SQL warehouse connection](/data-analysis/environment/fabrics/prophecy-fabrics). Otherwise, use the Table gem to [read from](/data-analysis/gems/source-target/table/bigquery-read) and [write to](/data-analysis/gems/source-target/table/bigquery-write) BigQuery. ## Create a BigQuery gem To create a BigQuery Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **BigQuery**. Then, click **Next**. In the **Location** tab, set your connection details and table location. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the table properties. To learn more, jump to [Source properties](#source-properties) and [Target properties](#target-properties). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source configuration Use these settings to configure a BigQuery Source gem for reading data. ### Source location | Parameter | Description | | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the source. For BigQuery tables, set to `bigquery`. | | Select or create connection | Select or create a new [BigQuery connection](/data-analysis/environment/connections/bigquery) in the Prophecy fabric you will use. | | Dataset | Dataset containing the table you want to read from. | | Name | Exact name of the BigQuery table to read data from. | ### Source properties Infer or manually configure the schema of your Source gem. Optionally, add a description for your table. Additional properties are not supported at this time. ## Target configuration Use these settings to configure a BigQuery Target gem for writing data. ### Target location | Parameter | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | | Format type | Table format for the target. For BigQuery tables, set to `bigquery`. | | Select or create connection | Choose or create a [BigQuery connection](/data-analysis/environment/connections/bigquery) in the Prophecy fabric you will use. | | Dataset | Dataset where the target table will be created or updated. | | Name | Name of the BigQuery table to write data to. If the table doesn't exist, it will be created automatically. | ### Target properties | Property | Description | Default | | ----------- | --------------------------------------------------------------------------------------------------------------- | ------- | | Description | Description of the table. | None | | Write Mode | Whether to overwrite the table completely, append new data to the table, or throw an error if the table exists. | None | ## Cross-workspace access If your fabric uses BigQuery as the SQL warehouse, you can't select BigQuery in an external Source or Target gem. Instead, you must use Table gems, which are limited to the BigQuery warehouse defined in the SQL warehouse connection. To work with tables from a different BigQuery workspace, use [BigQuery sharing](https://cloud.google.com/bigquery/docs/analytics-hub-introduction). This lets you access shared resources without creating additional BigQuery connections. Prophecy implements this guardrail to avoid using external connections when the data can be made available in your warehouse. External connections introduce an extra data transfer step, which slows down pipeline execution and adds unnecessary complexity. For best performance, Prophecy always prefers reading and writing directly within the warehouse. # Databricks external table gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/databricks Read and write catalog tables in Databricks This gem runs in . ## Overview This page describes how to use Databricks external Source and Target gems to read from or write to tables. Only use an external Source and Target gem when Databricks is not the configured [SQL warehouse connection](/data-analysis/environment/fabrics/prophecy-fabrics). Otherwise, use the Table gem to [read from](/data-analysis/gems/source-target/table/bigquery-read) and [write to](/data-analysis/gems/source-target/table/bigquery-write) Databricks. If you're working with file types like CSV or Parquet from Databricks file storage, see [File types](/data-analysis/gems/source-target/file) for guidance. This page focuses only on catalog tables. ## Create a Databricks gem To create a Databricks Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **Databricks** under **Table**. Do not select Databricks under **File**. Then, click **Next**. In the **Location** tab, set your connection details and table location. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the table properties. To learn more, jump to [Source properties](#source-properties) and [Target properties](#target-properties). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source configuration Use these settings to configure a Databricks Source gem for reading data. ### Source location | Parameter | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the source. For Databricks tables, set to `databricks`. | | Select or create connection | Select or create a new [Databricks connection](/data-analysis/environment/connections/databricks) in the Prophecy fabric you will use. | | Database | Database including the schema where the table is located. | | Schema | Schema containing the table you want to read from. | | Name | Exact name of the Databricks table to read data from. | ## Target configuration Use these settings to configure a Databricks Target gem for writing data. ### Target location | Parameter | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the target. For Databricks tables, set to `databricks`. | | Select or create connection | Select or create a new [Databricks connection](/data-analysis/environment/connections/databricks) in the Prophecy fabric you will use. | | Database | Database including the schema where the table is/will be located. | | Schema | Schema where the target table will be created or updated. | | Name | Name of the Databricks table to write data to. If the table doesn't exist, it will be created automatically. | ### Target properties | Property | Description | Default | | ----------- | --------------------------------------------------------------------------------------------------------------- | ------- | | Description | Description of the table. | None | | Write Mode | Whether to overwrite the table completely, append new data to the table, or throw an error if the table exists. | None | ## Cross-workspace access If your fabric uses Databricks as the SQL warehouse, you can't select Databricks in an external Source or Target gem. Instead, you must use Table gems, which are limited to the Databricks warehouse defined in the SQL warehouse connection. To work with tables from a different Databricks workspace, use [Delta Sharing](https://docs.databricks.com/aws/en/delta-sharing/). Delta Sharing lets you access data across workspaces without creating additional Databricks connections. Prophecy implements this guardrail to avoid using external connections when the data can be made available in your warehouse. External connections introduce an extra data transfer step, which slows down pipeline execution and adds unnecessary complexity. For best performance, Prophecy always prefers reading and writing directly within the warehouse. # SAP HANA external table gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/hana/hana Read and write from SAP HANA This gem runs in . ## Overview The SAP HANA Source and Target gems let you connect Prophecy pipelines to SAP HANA tables for reading and writing data. This page outlines how to configure SAP HANA sources and targets using the appropriate connections, locations, and properties. ## Create a SAP HANA gem To create a SAP HANA Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **Hana**. Then, click **Next**. In the **Location** tab, set your connection details and table location. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the table properties. To learn more, jump to [Source properties](#source-properties) and [Target properties](#target-properties). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source configuration Use these settings to configure a SAP HANA Source gem for reading data. ### Source location | Parameter | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the source. For SAP HANA tables, set to `hana`. | | Select or create connection | Select or create a new [SAP HANA connection](/data-analysis/environment/connections/hana) in the Prophecy fabric you will use. | | Read Using | How to define the table location.
  • **Table**: Provide the schema and table name.
  • **Query**: Select the table using a SQL query.
| | Schema (Table only) | Schema in SAP HANA where the table is located. | | Name (Table only) | Name of the SAP HANA table to read data from. | | Query (Query only) | SQL query used to retrieve a table. | ### Source properties Infer or manually configure the schema of your Source gem. Optionally, add a description for your table. ## Target configuration Use these settings to configure a SAP HANA Target gem for writing data. ### Target location | Parameter | Description | | --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the source. For SAP HANA tables, set to `Hana`. | | Select or create connection | Select or create a new [SAP HANA connection](/data-analysis/environment/connections/hana) in the Prophecy fabric you will use. | | Schema | Schema in SAP HANA where the target table resides or will be created. Select the **+** sign on the field to use a configuration variable or secret in your schema definition. | | Name | Name of the SAP HANA table to write data to. Select the **+** sign on the field to use a configuration variable or secret in your table definition. If the table doesn't exist, it will be created automatically. | ### Target properties Review schema of your Target gem and optionally update the metadata of columns in the schema. You can also add a description for your table. #### Generated columns Generated columns are specific to SAP HANA and auto-generate values for new rows. You can define generated columns in the Properties tab via column-level metadata. Learn more in [Generated columns](/data-analysis/gems/source-target/external-table/hana/identity-columns). ### Write options Control how data is written into the target table during each run of the pipeline. Choose whether to overwrite, append, or merge rows. | Mode | Description | | ---------------------------------------------- | -------------------------------------------------------------------------------- | | Wipe & Replace Table | Deletes the existing table and creates a new one with the incoming data. | | Write Fresh Table; Error if Exists | Creates a new table, and fails if a table with the same name already exists. | | Append Row | Inserts all incoming rows to the existing table without modifying existing data. | | Merge - Upsert Row | Updates if keys match, insert if not. | | Merge - SCD 2 | Stores and manages the current and historical data over time. | | Merge - Update Row | Only updates columns where key matches. | | Merge - Delete Row | Deletes rows where key matches. | | Merge - Delete Row if Exists; Otherwise Insert | Deletes rows where key matches, otherwise insert a new row. | | Merge - Delete Row if Exists; then Insert | Deletes rows where key matches and always inserts new rows. | #### Advanced options Configure the following for **Wipe & Replace Table**, **Write Fresh Table; Error if exists**, and **Append Row** modes. | Parameter | Description | | --------------- | -------------------------------------------------------------------------- | | Parallelism | The number of insert queries to execute in parallel. | | Rows Per Insert | The number of rows included in each insert query to the SAP HANA database. | #### Merge parameters Configure the following for any **Merge**. | Parameter | Required For | Description | | --------------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | | Merge condition | All merge modes | Specify one or more key columns that determine how incoming rows match existing rows in the target table. These columns act as the unique identifiers. | | Merge columns | Upsert row, Update row, Delete row if exists; otherwise insert | Select the specific columns to update or write when a key match is found. This controls which fields are modified during updates or inserts. | ### Example: Merge - Upsert Row This example demonstrates how the **Upsert row** merge mode works when updating an SAP HANA table. #### Existing table in SAP HANA Your SAP HANA table currently contains the following product inventory data:
| product\_id | category | price | stock\_quantity | last\_restocked | | ----------- | ----------- | ------ | --------------- | --------------- | | 2001 | Electronics | 99.99 | 100 | 2025-07-15 | | 2002 | Apparel | 29.99 | 50 | 2025-07-10 | | 2003 | Electronics | 149.99 | 20 | 2025-07-20 |
#### Merge configuration You configure the merge using the following settings: * **Merge Mode:** Upsert Row (update if key matches, insert if not) * **Unique Key:** `product_id` * **Merge Columns:** `stock_quantity`, `last_restocked` (only these columns are updated/inserted) #### Incoming data New inventory updates are coming through your pipeline:
| product\_id | category | stock\_quantity | last\_restocked | | ----------- | ----------- | --------------- | --------------- | | 2001 | Electronics | 150 | 2025-08-01 | | 2003 | Gaming | 0 | 2025-07-30 | | 2005 | Apparel | 85 | 2025-08-02 |
#### Result The Target gem processes each incoming row by checking if the `product_id` already exists in the target table: After the Target gem completes, your table looks like this:
| product\_id | category | price | stock\_quantity | last\_restocked | | ----------- | ----------- | ------ | --------------- | --------------- | | 2001 | Electronics | 99.99 | 150 | 2025-08-01 | | 2002 | Apparel | 29.99 | 50 | 2025-07-10 | | 2003 | Electronics | 149.99 | 0 | 2025-07-30 | | 2005 | null | null | 85 | 2025-08-02 |
##### Product 2001 * The key matches an existing row, so the system updates the record. * `stock_quantity` changes from **100 β†’ 150**, and `last_restocked` changes from **2025-07-15 β†’ 2025-08-01**. * `category` and `price` remain unchanged. ##### Product 2003 * The key matches an existing row, so the system updates the record. * `stock_quantity` changes from **20 β†’ 0**, and `last_restocked` changes from **2025-07-20 β†’ 2025-07-30**. * `category` and `price` remain unchanged. ##### Product 2005 * The key does not exist, so a new row is inserted. * Only the merge columns receive values: `stock_quantity` is set to **85** and `last_restocked` is set to **2025-08-02**. * `category` and `price` are set to **null** because they are not in the specified merge columns. # Generated columns Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/hana/identity-columns Define generated columns in HANA target gems Generated columns automatically generate values for inserted rows. Configure the generation behavior based on your use case. ## Generated column types The following generated column types are supported for SAP HANA: * `GENERATED ALWAYS AS IDENTITY` * `GENERATED BY DEFAULT AS IDENTITY` * `GENERATED ALWAYS AS (expression)` ### GENERATED ALWAYS AS IDENTITY When a column has the `GENERATED ALWAYS AS IDENTITY` type, it: * Accepts sequence parameters such as `START WITH` and `INCREMENT BY` to control how values are generated. * Generates values automatically based solely on the sequence parameters. * Creates an IDENTITY column managed entirely by SAP HANA. * Ignores values provided during an insert. **Example**: Assume you need a column that auto-increments primary keys for users or products. You never want users or inserts to override the value. * Column: `user_id` * Type: `GENERATED ALWAYS AS IDENTITY` * Always As: `START WITH 1 INCREMENT BY 1` Here, `user_id` will automatically start at 1 and increment by 1 for every new row. Attempts to insert values will be ignored. ### GENERATED BY DEFAULT AS IDENTITY When a column has the `GENERATED BY DEFAULT AS IDENTITY` type, it: * Accepts sequence parameters such as `START WITH` and `INCREMENT BY` to control how values are generated. * Generates values automatically *when no value is provided in an insert*. * Creates an IDENTITY column that can accept manual values if they are included in the insert. * Adjusts the internal sequence automatically if a manually provided value is higher (or lower, in the case of a negative increment) than the current sequence. **Example**: Assume you need to migrate legacy data while keeping original IDs, but still allowing auto-increment for new rows. * Column: `customer_id` * Type: `GENERATED BY DEFAULT AS IDENTITY` * Always As: `START WITH 1000 INCREMENT BY 1` Suppose you have existing customers with IDs 1000–1050. When inserting these rows, you can provide the IDs explicitly. New customers inserted without an ID will automatically get the next sequence value (1051, 1052, ...) by default. For more information on sequence parameters, see [CREATE SEQUENCE Statement](https://help.sap.com/docs/SAP_HANA_PLATFORM/4fe29514fd584807ac9f2a04f6754767/20d509277519101489029c064d468c5d.html). This includes reference information on syntax and default values. ### GENERATED ALWAYS AS (expression) When a column has the `GENERATED ALWAYS AS (expression)` type, it: * Computes values using a defined expression based on other columns in the table. * Updates automatically whenever referenced columns change. * Cannot be manually overridden; all values are derived from the expression. **Example**: Assume you want a column that always concatenates the first and last name columns. * Column: `full_name` * Type: `GENERATED ALWAYS AS` * Expression: `first_name || ' ' || last_name` This ensures that `full_name` will always reflect the `first_name` and `last_name`. It cannot be manually changed. ## Define generated columns To define a column as an IDENTITY column in a new target table: 1. Add a new Target gem to your pipeline canvas. 2. In the **Type** tab, select HANA as the target type. 3. In the **Location** tab, configure the location where the new table is written. 4. In the **Properties** tab, review the schema of the target table. 5. To define a column as an IDENTITY column: * Click on the dropdown arrow that appears on column hover. This expands the column metadata options. * Select a generated column type from the **Type** dropdown. * Provide a sequence parameter or expression in the **Always As** field. 6. Click **Save** to save your changes to the gem. When you run the gem, Prophecy recognizes the column as generated and lets SAP HANA handle value generation automatically. If your pipeline includes a field matching a `GENERATED ALWAYS AS IDENTITY` column, Prophecy automatically excludes it from insert operations to avoid conflicts with SAP HANA's internal sequence generation. For `GENERATED BY DEFAULT AS IDENTITY` columns, Prophecy can insert values when they are explicitly provided in your dataset. Otherwise, SAP HANA automatically generates the value. Hana table schema ## Write to tables with generated columns If you write to a HANA table that already contains an IDENTITY column, the HANA Gem: * Automatically detects the column in your target table * Lets SAP HANA handle the auto-increment process Prophecy does not label generated columns in metadata when the table originates from SAP HANA. This does not affect write behavior, but you might not catch that the column is generated. ## What's next For detailed information on functionality and constraints, see the `` and `` sections of the [CREATE TABLE Statement](https://help.sap.com/docs/SAP_HANA_PLATFORM/4fe29514fd584807ac9f2a04f6754767/20d58a5f75191014b2fe92141b7df228.html) page in the SAP HANA documentation. # MongoDB external table gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/mongodb Read and write from MongoDB This gem runs in . ## Overview This page describes how to configure MongoDB Source and Target gems, including connection setup, schema options, and available write modes. Use the MongoDB Source or Target gem to read from or write to MongoDB collections within your pipeline. ## Create a MongoDB gem To create a MongoDB Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **MongoDB**. Then, click **Next**. In the **Location** tab, set your connection details and collection location. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the table properties. To learn more, jump to [Source properties](#source-properties) and [Target properties](#target-properties). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source configuration Use these settings to configure a MongoDB Source gem for reading data from a collection. ### Source location | Parameter | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the target. For MongoDB, set to `mongodb`. | | Select or create connection | Select or create a new [MongoDB connection](/data-analysis/environment/connections/mongodb) in the Prophecy fabric you will use. | | Database | Database containing the table you want to read from. | | Name | Name of the MongoDB table to read. | ### Source properties | Property | Description | Default | | -------------------------------------------- | ---------------------------------------------------------------------- | ------- | | Description | Description of the table. | None | | No. of docs to consider for Schema inference | Number of documents to sample from the collection to infer the schema. | None | ## Target configuration Use these settings to configure a MongoDB Target gem for writing data to a collection. ### Target location | Parameter | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the target. For MongoDB, set to `mongodb`. | | Select or create connection | Select or create a new [MongoDB connection](/data-analysis/environment/connections/mongodb) in the Prophecy fabric you will use. | | Database | Database where the target table will be created or updated. | | Name | Name of the MongoDB table to write data to. If the table doesn't exist, it will be created automatically. | ### Target properties | Property | Description | Default | | ----------- | ---------------------------------------------------------------------------------------------------- | ------- | | Description | Description of the table. | None | | Write Mode | Whether to overwrite the table, append new data to the table, or throw an error if the table exists. | None | # MSSQL external table gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/mssql Read and write from MSSQL database This gem runs in . ## Overview This page describes how to configure Microsoft SQL Server (MSSQL) Source and Target gems, including connection setup, schema options, and available write modes. Use the MSSQL Source or Target gem to read from or write to the SQL server within your pipeline. ## Create an MSSQL gem To create an MSSQL Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **MSSQL**. Then, click **Next**. In the **Location** tab, set your connection details and table location. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the table properties. To learn more, jump to [Source properties](#source-properties) and [Target properties](#target-properties). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source configuration Use these settings to configure an MSSQL Source gem for reading data. ### Source location | Parameter | | Description | | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ----------- | | Format type | Table format for the source. For MSSQL tables, set to `mssql`. | | | Select or create connection | Select or create a new [MSSQL connection](/data-analysis/environment/connections/mssql) in the Prophecy fabric you will use. | | | Database | Database containing the table you want to read from. | | | Schema | Schema within the database where the table is located. | | | Name | Exact name of the MSSQL table to read data from. | | ### Source properties Infer or manually configure the schema of your Source gem. Optionally, add a description for your table. Additional properties are not supported at this time. ## Target configuration Use these settings to configure an MSSQL Target gem for writing data. ### Target location | Parameter | | Description | | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ----------- | | Format type | Table format for the source. For MSSQL tables, set to `mssql`. | | | Select or create connection | Select or create a new [MSSQL connection](/data-analysis/environment/connections/mssql) in the Prophecy fabric you will use. | | | Database | Database where the target table will be created or updated. | | | Schema | Schema within the database where the target table resides or will be created. | | | Name | Name of the MSSQL table to write data to. If the table doesn't exist, it will be created automatically. | | ### Target properties | Property | Description | Default | | ----------- | ---------------------------------------------------------------------------------------------------- | ------- | | Description | Description of the table. | None | | Write Mode | Whether to overwrite the table, append new data to the table, or throw an error if the table exists. | None | # Oracle external table gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/oracle Read and write from Oracle This gem runs in . ## Overview This page describes how to configure Oracle Source gems, including connection setup and schema options. Note that Prophecy does not support Oracle as a target destination. ## Create an Oracle gem To create an Oracle Source gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **Oracle**. Then, click **Next**. In the **Location** tab, set your connection details and table location. To learn more, jump to [Source location](#source-location). In the **Properties** tab, set the table properties. To learn more, jump to [Source properties](#source-properties). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source configuration Use these settings to configure an Oracle Source gem for reading data. ### Source location | Parameter | Description | | --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the source. For Oracle tables, set to `oracle`. | | Select or create connection | Select or create a new [Oracle connection](/data-analysis/environment/connections/oracle) in the Prophecy fabric you will use. | | Read using | Choose table or query.
  • **Table**: Provide the schema and name of the table you want to read.
  • **Query**: Enter a SQL query directly in the gem to select a table.
| ### Source properties Infer or manually configure the schema of your Source gem. Optionally, add a description for your table. Additional properties are not supported at this time. # Postgres external table gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/postgres Read and write from PostgreSQL database This gem runs in . This page describes how to configure PostgreSQL Source and Target gems, including connection setup, schema options, and available write modes. Use the Postgres Source or Target gem to read from or write to a PostgreSQL database within your pipeline. ## Create a Postgres gem To create a Postgres Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **Postgres**. Then, click **Next**. In the **Location** tab, set your connection details and table location. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the table properties. To learn more, jump to [Source properties](#source-properties) and [Target properties](#target-properties). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source configuration Use these settings to configure a Postgres Source gem for reading data. ### Source location | Parameter | Description | | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the source. For PostgreSQL tables, set to `postgres`. | | Select or create connection | Select or create a new [Postgres connection](/data-analysis/environment/connections/postgres) in the Prophecy fabric you will use. | | Schema | Schema containing the table you want to read from. | | Name | Exact name of the PostgreSQL table to read data from. | | Query | SQL query to read data from. If provided, this takes precedence over the schema and table name fields. | ### Source properties Infer or manually configure the schema of your Source gem. Optionally, add a description for your table. Additional properties are not supported at this time. ## Target configuration Use these settings to configure a Postgres Target gem for writing data. ### Target location | Parameter | Description | | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the target. For PostgreSQL tables, set to `postgres`. | | Select or create connection | Select or create a new [Postgres connection](/data-analysis/environment/connections/postgres) in the Prophecy fabric you will use. | | Schema | Schema where the target table resides or will be created. | | Name | Name of the PostgreSQL table to write data to. If the table doesn't exist, it will be created automatically. | ### Target properties | Property | Description | Default | | ----------- | ---------------------------------------------------------------------------------------------------- | ------- | | Description | Description of the table. | None | | Write Mode | Whether to overwrite the table, append new data to the table, or throw an error if the table exists. | None | ## Configure a Postgres connection Use these settings to create a Postgres connection: | Parameter | Description | | ----------------------- | ------------------------------------------------------------------------- | | Connection Name | Name for the connection. | | Server | PostgreSQL server hostname or IP address. | | Port | Port used to connect to the PostgreSQL server. | | Database | PostgreSQL database name. | | Username | Username used for pipeline development and scheduled execution. | | Password | Password used for pipeline development and scheduled execution. | | Knowledge Graph Indexer | Choose whether to enable the Knowledge Graph Indexer for this connection. | # Amazon Redshift external table gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/redshift Read and write from Redshift This gem runs in . ## Overview The Redshift Source and Target gems let you connect Prophecy pipelines to Amazon Redshift tables for reading and writing data. This page outlines how to configure Redshift sources and targets using the appropriate connections, locations, and properties. ## Create a Redshift gem To create a Redshift Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **Redshift**. Then, click **Next**. In the **Location** tab, set your connection details and table location. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the table properties. To learn more, jump to [Source properties](#source-properties) and [Target properties](#target-properties). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source configuration Use these settings to configure a Redshift Source gem for reading data. ### Source location | Parameter | Description | | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the source. For Amazon Redshift tables, set to `redshift`. | | Select or create connection | Select or create a new [Redshift connection](/data-analysis/environment/connections/redshift) in the Prophecy fabric you will use. | | Database | Database containing the table you want to read from. | | Schema | Schema within the database where the table is located. | | Name | Exact name of the Amazon Redshift table to read data from. | ### Source properties Infer or manually configure the schema of your Source gem. Optionally, add a description for your table. Additional properties are not supported at this time. ## Target configuration Use these settings to configure a Redshift Target gem for writing data. ### Target location | Parameter | Description | | --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the source. For Amazon Redshift tables, set to `redshift`. | | Select or create connection | Select or create a new [Amazon Redshift connection](/data-analysis/environment/connections/redshift) in the Prophecy fabric you will use. | | Database | Database where the target table will be created or updated. | | Schema | Schema within the database where the target table resides or will be created. | | Name | Name of the Amazon Redshift table to write data to. If the table doesn't exist, it will be created automatically. | ### Target properties | Property | Description | Default | | ----------- | ---------------------------------------------------------------------------------------------------- | ------- | | Description | Description of the table. | None | | Write Mode | Whether to overwrite the table, append new data to the table, or throw an error if the table exists. | None | # Salesforce external table gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/salesforce Read and write from Salesforce This gem runs in . ## Overview The Salesforce gem enables you to read data from Salesforce directly into your Prophecy pipelines. The gem supports both SOQL (Salesforce Object Query Language) for querying standard Salesforce objects and SAQL (Salesforce Analytics Query Language) for accessing analytical datasets. ## Create a Salesforce gem To create a Salesforce Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **Salesforce**. Then, click **Next**. In the **Location** tab, set your connection details and query configuration. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the table properties. To learn more, jump to [Source properties](#source-properties) and [Target properties](#target-properties). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source configuration Use these settings to configure a Salesforce Source gem for reading data. ### Source location | Parameter | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the source. For Salesforce tables, set to `salesforce`. | | Select or create connection | Select or create a new [Salesforce connection](/data-analysis/environment/connections/salesforce) in the Prophecy fabric you will use. | | Query Mode | Specify the table you would like to read using a query. Learn more in [Query modes](#query-modes). | ### Query modes You have three options for retrieving data from Salesforce. #### SAQL Use a [SAQL](https://developer.salesforce.com/docs/atlas.en-us.bi_dev_guide_saql.meta/bi_dev_guide_saql/bi_saql_intro.htm) query to access CRM Analytics datasets (formerly known as Wave Analytics). For example: ``` q = load "Account"; q = foreach q generate 'Id' as 'Id', 'Name' as 'Name'; q = limit q 5; ``` #### SOQL Use a [SOQL](https://developer.salesforce.com/docs/atlas.en-us.soql_sosl.meta/soql_sosl/sforce_api_calls_soql.htm) query to read structured data from Salesforce objects such as `Account` or `Contact` tables. For example: ```SQL theme={null} SELECT Id, Name, Industry, CreatedDate FROM Account WHERE CreatedDate >= LAST_N_DAYS:30 ``` When using the `FIELDS(ALL)` keyword, the response is [limited to 200 rows per call](https://developer.salesforce.com/docs/atlas.en-us.soql_sosl.meta/soql_sosl/sforce_api_calls_soql_select_fields.htm#limiting_result_rows). Prophecy does not support SOQL parent-to-child relationship queries that use nested subqueries. Query parent and child objects separately, then join the datasets in Prophecy. For example, the following query is not supported: ```sql theme={null} SELECT Id, Name, (SELECT FirstName, LastName FROM Contacts) FROM Account ``` #### Salesforce Objects Retrieve all records from a specific Salesforce object by name, without writing a query. Use this option when the object exceeds the default 200 row limit of `FIELDS(ALL)`. ### Source properties The following properties are available for the Salesforce Source gem. These properties only apply to tables retrieved with the **SOQL** or **Salesforce Object** query mode. | Property | Description | | ------------------------------------- | -------------------------------------------------------------------------------------------------- | | Enable bulk query | Enable to run the query as a batch job in the background for better performance on large datasets. | | Retrieve deleted and archived records | Enable to include soft-deleted and archived records in the returned table. | ## Target configuration Use these settings to configure a Salesforce Target gem for writing data. The Salesforce Target gem only supports writing to Salesforce Objects. You cannot write to CRM Analytics datasets. ### Target location Use the following parameters to define the write location. | Parameter | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the source. For Salesforce tables, set to `salesforce`. | | Select or create connection | Select or create a new [Salesforce connection](/data-analysis/environment/connections/salesforce) in the Prophecy fabric you will use. | | Object Name | Name of the Salesforce Object to write to. If the object doesn't exist, it will be created automatically. | ### Target properties The following properties are available for the Salesforce Target gem. | Property | Description | Default | | ----------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | | Description | Text description of the target table. Use this field to document the purpose or content of the data being written. | None | | Write Mode | Defines how records are written to Salesforce.
  • Upsert: Update existing records if a match is found; otherwise, insert new records.
  • Insert: Add new records without modifying existing ones.
  • Update: Update existing records if a match is found.
  • Delete: Remove existing records if a match is found.
| Upsert | | External ID Field | Specifies the Salesforce field used as the unique key when performing Upsert or Update operations. The field must be defined as an External ID in Salesforce. | None | | Use Bulk API | Enable to run write operations using the Salesforce [Bulk API](https://developer.salesforce.com/docs/atlas.en-us.api_asynch.meta/api_asynch/bulk_api_2_0.htm). | Disabled | # Snowflake external table gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/snowflake Read and write from Snowflake This gem runs in . ## Overview The Snowflake Source and Target gems let you connect Prophecy pipelines to Snowflake tables for reading and writing data. This page outlines how to configure Snowflake sources and targets using the appropriate connections, locations, and properties. ## Create a Snowflake gem To create a Snowflake Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **Snowflake**. Then, click **Next**. In the **Location** tab, set your connection details and table location. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the table properties. To learn more, jump to [Source properties](#source-properties) and [Target properties](#target-properties). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source configuration Use these settings to configure a Snowflake Source gem for reading data. ### Source location | Parameter | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the source. For Snowflake tables, set to `snowflake`. | | Select or create connection | Choose or create a [Snowflake connection](/data-analysis/environment/connections/snowflake) in the Prophecy fabric you will use. | | Database | Snowflake database containing the table you want to read from. | | Schema | Schema within the database where the table is located. | | Name | Exact name of the Snowflake table to read data from. | ### Source properties Infer or manually configure the schema of your Source gem. Optionally, add a description for your table. Additional properties are not supported at this time. ## Target configuration Use these settings to configure a Snowflake Target gem for writing data. ### Target location | Parameter | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------- | | Format type | Table format for the target. For Snowflake tables, set to `snowflake`. | | Select or create connection | Choose or create a [Snowflake connection](/data-analysis/environment/connections/snowflake) in the Prophecy fabric you will use. | | Database | Snowflake database where the target table will be created or updated. | | Schema | Schema within the database where the target table resides or will be created. | | Name | Name of the Snowflake table to write data to. If the table doesn't exist, it will be created automatically. | ### Target properties | Property | Description | Default | | ----------- | --------------------------------------------------------------------------------------------------------------- | ------- | | Description | Description of the table. | None | | Write Mode | Whether to overwrite the table completely, append new data to the table, or throw an error if the table exists. | None | # MSSQL on Azure Synapse dedicated SQL pool Source: https://docs.prophecy.ai/data-analysis/gems/source-target/external-table/synapse Read and write from an Azure Synapse dedicated SQL pool This gem runs in . ## Overview This page describes how to configure Source gems to read from Microsoft SQL Server hosted on Azure Synapse dedicated SQL pools. ## Prerequisites * Run Prophecy version 4.1.3 or higher. ## Create a Synapse gem To create a Synapse Source gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **Synapse**. Then, click **Next**. In the **Location** tab, set your connection details and table location. To learn more, jump to [Source location](#source-location). In the **Properties** tab, set the table properties. To learn more, jump to [Source properties](#source-properties). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source configuration Use these settings to configure a Source gem for reading data from Azure Synapse. ### Source location | Parameter | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------ | | Format type | Table format for the source. For MSSQL tables on Azure Synapse, set to `mssql`. | | Select or create connection | Existing or new [Azure Synapse connection](/data-analysis/environment/connections/synapse) in the Prophecy fabric. | | Database | Database matching the database defined in the [connection](/data-analysis/environment/connections/synapse). | | Schema | Schema within the database where the table is located. | | Name | Name of the table to read data from. | ### Source properties Infer or manually configure the schema of the table. Optionally, add a description for your table. Additional properties are not supported at this time. # Azure Data Lake Storage (ADLS) Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/adls Use Azure Data Lake Storage as a file source or target in a gem This gem runs in . Use a Source and Target gem to read from or write to [Azure Data Lake Storage](https://learn.microsoft.com/en-us/azure/storage/blobs/data-lake-storage-introduction) (ADLS) locations in Prophecy pipelines. This page covers supported file formats, how to create the gem, and how to configure connection details and paths for both Source and Target gems. ## Supported file formats | Format | Read | Write | | ---------------------------------------------------------------------------- | ---- | ----- | | [CSV](/data-analysis/gems/source-target/file/file-types/csv) | βœ” | βœ” | | [Fixed width](/data-analysis/gems/source-target/file/file-types/fixed-width) | βœ” | | | [JSON](/data-analysis/gems/source-target/file/file-types/json) | βœ” | βœ” | | [Parquet](/data-analysis/gems/source-target/file/file-types/parquet) | βœ” | βœ” | | [XLSX](/data-analysis/gems/source-target/file/file-types/excel) | βœ” | βœ” | | [XML](/data-analysis/gems/source-target/file/file-types/xml) | βœ” | βœ” | ## Create an ADLS gem To create an ADLS Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **ADLS**. Then, click **Next**. In the **Location** tab, set your file format and connection details. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the file properties. These vary based on the file type that you are working with. See the list of properties per file type, such as [CSV](/data-analysis/gems/source-target/file/file-types/csv). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source location When setting up an ADLS Source gem, you need to decide how Prophecy should locate files at runtime. There are two modes: * **Filepath**: Always read from a specific file path. You can use wildcards in your path definition. If multiple files match, they are unioned into a single output table. * **Configuration**: Dynamically read files provided by a [file arrival/change trigger](/data-analysis/production/scheduling/triggers) in the pipeline's schedule. When new or updated files are detected in the monitored directory, the trigger starts the pipeline and passes those files to the ADLS Source gem, which unions them into a single output table. Use the table below to understand how to configure each option. | Parameter | Description | | ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Select or create connection | Select an existing ADLS connection or [create a new one](/data-analysis/environment/connections/adls). | | Choose Path or Configuration | Choose between the following options.
  • **Filepath**: Read one file from a specified file path.
  • **Configuration**: Dynamically read files provided by a [file arrival/change trigger](/data-analysis/production/scheduling/triggers).
| | File Path
*Filepath option only* | Path to the file in the ADLS container. Supports wildcards.
Example: `abfss://container-name@account-name.dfs.core.windows.net/temp/dir/*.csv` | | Examine | Automatically determine the file format, compression type (if any), properties, and schema of the file to read. | | Select Configuration
*Configuration option only* | File arrival/change trigger [configuration](/data-analysis/production/scheduling/triggers#trigger-configuration) that provides the added or modified files for that run. | | Include filename Column | Appends a column containing the source filename for each row in the output table. | | Delete files after successfully processed | Deletes objects after they are successfully read. | | Move files after successfully processed | Moves objects to a specified directory after they are successfully read. | | Format type | Type of file to read, such as csv or json. | | Compression | The compression type of the file to read.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | ### Configuration If you select **Configuration**, the gem will only run successfully during a **triggered pipeline run**. This is because the gem expects files from the trigger. In other cases, such as an interactive run or an API-triggered run, there will be no files to read. In these situations, you will encounter the following error: ``` Failed due to: Unable to detect modified files for provided File Trigger ``` For the same reason, you'll see an error if you try to infer the schema in the **Properties** tab or load a preview in the **Preview** tab of the Source gem. ## Target location When setting up an ADLS Target gem, you need to set the location and file type to correctly write the file. | Parameter | Description | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Select an existing ADLS connection or [create a new one](/data-analysis/environment/connections/adls). | | File Path | ADLS path where the output file will be written.
Example: `abfss://container-name@account-name.dfs.core.windows.net/data/orders.csv` | | Examine | If an existing file exists at the target path, click **Examine** to automatically determine the file format and compression type (if any). | | Encryption Algorithm | The encryption algorithm to use when writing the file.

Supported algorithms: `AES-192`, `AES-296`, `BlowFish` | | Format type | Type of file to write, such as csv or json. | | Compression | The compression type to use when writing the file.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | Configure encryption in a Target gem to encrypt an entire file. Use the [DataEncoderDecoder](/data-analysis/gems/transform/encoder-decoder) gem to encrypt individual columns. # Databricks Volumes Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/databricks-volumes Use Databricks Volumes as a file source or target in a gem This gem runs in . ## Overview Use a Source or Target gem to read from or write to Databricks Volumes in Prophecy pipelines. This page covers supported file formats, how to create the gem, and how to configure connection details and paths for both Source and Target gems. ## Supported file formats | Format | Read | Write | | ---------------------------------------------------------------------------- | ---- | ----- | | [CSV](/data-analysis/gems/source-target/file/file-types/csv) | βœ” | βœ” | | [Fixed width](/data-analysis/gems/source-target/file/file-types/fixed-width) | βœ” | | | [JSON](/data-analysis/gems/source-target/file/file-types/json) | βœ” | βœ” | | [Parquet](/data-analysis/gems/source-target/file/file-types/parquet) | βœ” | βœ” | | [XLSX](/data-analysis/gems/source-target/file/file-types/excel) | βœ” | βœ” | | [XML](/data-analysis/gems/source-target/file/file-types/xml) | βœ” | βœ” | ## Create a Databricks Volumes gem To create a Databricks Volumes Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **Databricks** under **File**. Do not select Databricks under **Table**. Then, click **Next**. In the **Location** tab, set your file format and connection details. To learn more, jump to [Location](#location). In the **Properties** tab, set the file properties. These vary based on the file type that you are working with. See the list of properties per file type, such as [CSV](/data-analysis/gems/source-target/file/file-types/csv). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source location | Parameter | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Select an existing Databricks Volumes connection or [create a new one](/data-analysis/environment/connections/databricks). | | File Path | Path to the file in Databricks Volumes.
Example: `/Volumes/catalog/schema/volume/file.csv` | | Examine | Automatically determine the file format, compression type (if any), properties, and schema of the file to read. | | Format type | Type of file to read, such as `csv` or `json`. | | Compression | The compression type of the file to read.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | ## Target location | Parameter | Description | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Select an existing Databricks Volumes connection or [create a new one](/data-analysis/environment/connections/databricks). | | File Path | Path to the file in Databricks Volumes.
Example: `/Volumes/catalog/schema/volume/file.csv` | | Examine | If an existing file exists at the target path, click **Examine** to automatically determine the file format and compression type (if any). | | Encryption Algorithm | The encryption algorithm to use when writing the file.

Supported algorithms: `AES-192`, `AES-296`, `BlowFish` | | Format type | Type of file to write, such as `csv` or `json`. | | Compression | The compression type to use when writing the file.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | Configure encryption in a Target gem to encrypt an entire file. Use the [DataEncoderDecoder](/data-analysis/gems/transform/encoder-decoder) gem to encrypt individual columns. # CSV file gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/file-types/csv Read and write CSV files This page describes the **CSV-specific properties** that appear in the **Properties** tab of Source and Target gems. These settings are the same for CSV files regardless of which connection type is configured in the gem (for example, S3, SFTP, or SharePoint). If you need details on configuring a Source or Target gem end to end (including all tabs such as **Location**), see the documentation for the specific file storage connection. You can also use the [upload file](/data-analysis/gems/source-target/table/upload-files) feature to use CSV files. These will be stored in the SQL warehouse configured in your fabric. ## Properties ### Source properties The following properties are available for the CSV Source gem. | Property | Description | Default | | ----------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------- | | Description | Description of the table. | None | | Separator | Character used to separate values in the CSV file. | `,` | | Header | Whether the first row is the column header. | True | | Null Value | String that represents a null or missing value in the CSV. | None | | Comment Character | Character used to denote lines in the file that should be treated as comments. | None | | Inference Data Sampling Limit | Maximum number of rows to sample for inferring the schema. | `0` | | File Encoding | Character set used to decode the CSV file when reading. See [supported encodings](#supported-character-encodings). | `UTF-8` | ### Target properties The following properties are available for the CSV Target gem. | Property | Description | Default | | -------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------- | | Description | Description of the table. | None | | Separator | Character used to separate values in the CSV file. | `,` | | Header | Whether to make the first row the column header. | True | | Null Value | String that represents a null or missing value in the CSV. | None | | Use CRLF as line separator | If enabled, lines in the CSV will end with `\r\n` (Windows-style newlines). | None | | File Encoding | Character set used to encode the CSV file when writing. See [supported encodings](#supported-character-encodings). | `UTF-8` | * UTF-8 - UTF-16 - ISO-8859-1 - ISO-8859-2 - ISO-8859-3 - ISO-8859-4 - ISO-8859-5 - ISO-8859-6 - ISO-8859-7 - ISO-8859-8 - ISO-8859-9 - ISO-8859-10 - ISO-8859-13 - ISO-8859-14 - ISO-8859-15 - ISO-8859-16 - Windows-1250 - Windows-1251 - Windows-1252 - Windows-1253 - Windows-1254 - Windows-1255 - Windows-1256 - Windows-1257 - Windows-1258 - Windows-874 - CodePage437 - CodePage850 - CodePage852 - CodePage855 - CodePage858 - CodePage860 - CodePage862 - CodePage863 * CodePage865 - CodePage866 - Macintosh - MacintoshCyrillic - KOI8R - KOI8U - XUserDefined - ASCII # Excel Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/file-types/excel Read and write Excel files This page describes the **Excel-specific properties** that appear in the **Properties** tab of Source and Target gems. These settings are the same for Excel files regardless of which connection type is configured in the gem (for example, S3, SFTP, or SharePoint). If you need details on configuring a Source or Target gem end to end (including all tabs such as **Location**), see the documentation for the specific file storage connection. You can also use the [upload file](/data-analysis/gems/source-target/table/upload-files) feature to use Excel files. These will be stored in the SQL warehouse configured in your fabric. ## Properties ### Source properties The following properties are available for the Excel Source gem. | Property | Description | Default | | ----------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | ---------------------- | | Description | Description of the table. | None | | Header | Whether the first row is the column header. | True | | Allow Undefined Rows | Whether to permit rows with all values undefined (null or empty). | True | | Allow Incomplete Rows | Whether to permit rows with missing values for some columns. | True | | Ignore Cell Formatting | Whether to apply the number format for the cell value or get the raw value. | True | | Sheet Reading method | Whether to read one sheet of data or union the data from multiple sheets into one table. Learn more in [Reading sheets](#reading-sheets). | Read single sheet data | | Skip Undefined Rows | Whether to skip rows where all values are undefined. | False | | Date Format Reference | Date format to use when parsing date values. | `2006-01-02` | | Time Format Reference | Time format to use when parsing time values. | `15:04:05` | | Timestamp Format Reference | Timestamp format to use when parsing date-time values. | `2006-01-02 15:04:05` | | Inference Data Sampling Limit | Maximum number of rows to sample for inferring the schema. | `0` | | Password | Password for password-protected sheets. | None | ### Reading sheets Depending on the **Sheet Reading method** you choose, you will need to provide additional details. #### Read single sheet data For this option, Prophecy only reads one sheet of the Excel file. | Additional property | Description | | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | | Sheet Name | Provide the name of the sheet to read. If the name you provide does not match the name of an existing sheet in the file, the gem will fail to run. | #### Read union of multiple sheet data For this option, Prophecy reads the data from each of the sheets that you specify. Then, the data is unioned into one output table. | Additional property | Description | | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Sheet Filter Condition | Choose how you define the set of sheets to read.
  • Sheet name prefix equals: Read sheets whose names start with the provided value.
  • Sheet name suffix equals: Read sheets whose names end with the provided value.
  • Sheet name contains value: Read sheets whose names contain the provided substring.
  • Sheet name is in below list (comma separated): Read sheets whose names exactly match any comma-separated value provided in Filter Value.
| | Filter Value | Provide the value used to evaluate the filter condition. | | Output Sheet Column Name | Provide the name of the column to append to the output table that contains the original sheet name for each row. | The union operation will only succeed if each sheet has the same schema. ## Target properties The following properties are available for the Excel Target gem. | Property | Description | Default | | ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------- | | Description | Description of the table. | None | | Sheet Write method | Choose how rows are written.
  • Write data to a single sheet: All rows are written to one sheet specified in Sheet Name/Sheet Column Name. The sheet is created if it does not exist.
  • Dynamic data write to multiple sheets: Rows are partitioned by the value in the column specified in Sheet Name/Sheet Column Name, and each partition is written to a sheet with that name. Sheets are created as needed.
| Write data to a single sheet | | Sheet Name/Sheet Column Name | Configure the destination name reference.
  • Single sheet: Provide the exact sheet name to write to (for example, Sheet1).
  • Multiple sheets: Provide the column that contains the target sheet name for each row.
| `Sheet1` | | Header | Whether to make the first row the column header. | True | | Ignore Cell Formatting | Whether to apply the number format for the cell value or get the raw value. | True | | Password | Password for password-protected sheets. | None | # Fixed-width Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/file-types/fixed-width Read fixed-width files This page describes the **Fixed-width-specific properties** that appear in the **Properties** tab of Source gems. These settings are the same for fixed-width files regardless of which connection type is configured in the gem (for example, S3, SFTP, or SharePoint). Writing fixed-width files using a Target gem is not supported. If you need details on configuring a Source gem end to end (including all tabs such as **Location**), see the documentation for the specific file storage connection. ## Source schema Define the schema of your dataset in the **Properties** tab of the gem. This determines the structure of your dataset. You must define the schema manually for fixed-width files. Schema inference isn't supported. Each row in the **Schema** table corresponds to a fixed-width column in the file and includes the following attributes: | Field | Description | | -------- | ----------------------------------------------------------------------------------------- | | Name | Name of the column. | | Type | Type of data in the column. | | Offset | Starting character position in the row for this column. | | Length | Number of characters for this column. | | Metadata | Additional information about the column. Some data types require certain metadata fields. | ### Metadata fields Hover over a row in the Schema table to reveal a dropdown arrow next to the Metadata field. Use this dropdown to add column details, such as a description or tags. Metadata is optional for the majority of data types but is sometimes required. The tables below list metadata fields specific to certain data types per SQL warehouse provider: | Data type | Specific metadata fields | Required | | --------- | ------------------------------------------------------------------------------------------------------ | -------- | | Array | **Number of occurrences**: Number of elements in the array. | Yes | | Decimal | **Scale**: Number of digits after the decimal point. | No | | Date | **Format**: Expected format of the date that can override the global default at the column level. | No | | Time | **Format**: Expected format of the time that can override the global default at the column level. | No | | Timestamp | **Format**: Expected format of the timestamp that can override the global default at the column level. | No | | Data type | Specific metadata fields | Required | | ---------- | ------------------------------------------------------------------------------------------------------ | -------- | | Array | **Number of occurrences**: Number of elements in the array. | Yes | | BigNumeric | **Scale**: Number of digits after the decimal point. | No | | Numeric | **Scale**: Number of digits after the decimal point. | No | | Datetime | **Format**: Expected format of the datetime that can override the global default at the column level. | No | | Time | **Format**: Expected format of the time that can override the global default at the column level. | No | | Date | **Format**: Expected format of the date that can override the global default at the column level. | No | | Timestamp | **Format**: Expected format of the timestamp that can override the global default at the column level. | No | | Data type | Specific metadata fields | Required | | --------- | ------------------------------------------------------------------------------------------------------ | -------- | | Array | **Number of occurrences**: Number of elements in the array. | Yes | | Date | **Format**: Expected format of the date that can override the global default at the column level. | No | | Decimal | **Scale**: Number of digits after the decimal point. | No | | Time | **Format**: Expected format of the time that can override the global default at the column level. | No | | Timestamp | **Format**: Expected format of the timestamp that can override the global default at the column level. | No | For a complete list of supported data types, visit [Supported data types](/data-analysis/gems/data-types). ## Source properties The following properties are available to customize how fixed-width files are read. | Property | Description | Default | | ------------------------------------ | ------------------------------------------------------------------------------- | --------------------- | | Description | Description of the table. | None | | Line Delimited | Whether each record ends with a newline character. | Disabled | | Strip Trailing Blanks | Strip trailing whitespace from string values when reading the data. | Disabled | | Number of Initial Rows to Skip | Number of lines to skip before the data begins, usually `1` for the header row. | `0` | | Number of Bytes to Skip Between Rows | Usually used by the Transpiler. | `0` | | Date Format Reference | Global default format for parsing date columns. | `2006-01-02` | | Time Format Reference | Global default format for parsing time columns. | `15:04:05` | | Timestamp Format Reference | Global default format for parsing timestamp columns. | `2006-01-02 15:04:05` | | Decimal Point | Character used as a decimal point. | `.` | # JSON file gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/file-types/json Read and write JSON files This page describes the **JSON-specific properties** that appear in the **Properties** tab of Source and Target gems. These settings are the same for JSON files regardless of which connection type is configured in the gem (for example, S3, SFTP, or SharePoint). If you need details on configuring a Source or Target gem end to end (including all tabs such as **Location**), see the documentation for the specific file storage connection. You can also use the [upload file](/data-analysis/gems/source-target/table/upload-files) feature to use JSON files. These will be stored in the SQL warehouse configured in your fabric. ## Properties ### Source properties The following properties are available for the JSON Source gem. | Property | Description | Default | | ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------- | | Schema | Define the structure of your data. This includes the column names and the data type for each column (such as String, Integer, etc.).

Click **Infer Schema** to automatically detect the structure from the file. | None | | Description | Add a description of the table.

Click **Auto-description** to automatically generate the description. | None | | Multiple documents per file | Select whether the file contains multiple JSON objects separated by newline. | False | | Inference Data Sampling Limit | Define the maximum number of rows to sample for inferring the schema. Set to `0` to use all rows. | `0` | ### Target properties The following properties are available for the JSON Target gem. | Property | Description | Default | | --------------------------- | ---------------------------------------------------------------------------------------------------------------- | ------- | | Schema | Review the structure of your data. Prophecy automatically detects the schema from the input gem. | None | | Description | Add a description of the table.

Click **Auto-description** to automatically generate the description. | None | | Multiple documents per file | Select whether the file will contain multiple JSON objects separated by newline. | False | | Enable indentation | Select whether to format the JSON output with indentation for readability. | False | | Indentation size | Define the number of spaces to use for each indentation level when indentation is enabled. | `2` | | Escape HTML characters | Select whether to escape HTML characters (such as `<`, `>`, and `&`) in the JSON output. | True | ## Schema validation Prophecy lets you enable schema validation for JSON files in a Source gem. Use schema validation to ensure JSON files conform to a predefined structure before ingesting the data. JSON schema validation works for all file storage connections. ### Prerequisites To use schema validation, you need a JSON Schema file in the same directory as your source file. If your schema uses `$ref` to reference external schemas, those referenced schema files must also be in the same directory. Learn how to [write JSON Schema files](https://json-schema.org/learn). ### Set up schema validation 1. In the Location tab of a Source gem, toggle **Enable JSON Schema Validation**. 2. Provide a path to the JSON file that you will use to validate against. 3. Open the Properties tab. 4. Click **Infer Schema**. If the source file schema matches the validation schema, schema inference will run successfully. ### Example The following example shows a source file and a corresponding validation file that would validate successfully. ```json users.json theme={null} { "id": 1, "name": "John Doe", "email": "john.doe@example.com" } ``` ```json users-schema.json theme={null} { "$schema": "http://json-schema.org/draft-07/schema#", "type": "object", "properties": { "id": { "type": "integer" }, "name": { "type": "string" }, "email": { "type": "string", "format": "email" } }, "required": ["id", "name", "email"] } ``` ### Troubleshooting If the schemas do not match: * Schema inference will fail. * The Source gem will fail to run. To troubleshoot, look for the error in the [runtime logs](/data-analysis/development/runs/runtime-logs). Here is an example error: ```text wrap theme={null} Failed due to error in "OrchestrationSource_0". Error: JSON validation failed against JSON schema for file /Volumes/pipelinehub/dev/json-schema-validate/users.json: 1.data: Additional property location is not allowed; 1.data: Additional property lastLogin is not allowed ``` # Parquet file gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/file-types/parquet Read and write Parquet files This page describes the **Parquet-specific properties** that appear in the **Properties** tab of Source and Target gems. These settings are the same for Parquet files regardless of which connection type is configured in the gem (for example, S3, SFTP, or Databricks). If you need details on configuring a Source or Target gem end to end (including all tabs such as **Location**), see the documentation for the specific file storage connection. ## Properties ### Source properties The following properties are available for the Parquet Source gem. | Property | Description | Default | | ----------------------------- | ---------------------------------------------------------- | ------- | | Description | Description of the table. | None | | Multiple documents per file | Whether the file contains multiple Parquet documents. | False | | Inference Data Sampling Limit | Maximum number of rows to sample for inferring the schema. | `0` | ### Target properties The following properties are available for the Parquet Target gem. | Property | Description | Default | | ----------- | ------------------------- | ------- | | Description | Description of the table. | None | # Text file gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/file-types/text Read and write text files This page describes the **Text-specific properties** that appear in the **Properties** tab of Source and Target gems. These settings are the same for text files regardless of which connection type is configured in the gem (for example, S3, SFTP, or SharePoint). If you need details on configuring a Source or Target gem end to end (including all tabs such as **Location**), see the documentation for the specific file storage connection. ## Properties ### Source properties The following properties are available for the Text Source gem. | Property | Description | Default | | ----------------------------- | ---------------------------------------------------------- | ------- | | Description | Description of the table. | None | | Inference Data Sampling Limit | Maximum number of rows to sample for inferring the schema. | `0` | ### Target properties The following properties are available for the Text Target gem. | Property | Description | Default | | ----------- | ------------------------- | ------- | | Description | Description of the table. | None | # XML file gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/file-types/xml Read and write XML files This page describes the **XML-specific properties** that appear in the **Properties** tab of Source and Target gems. These settings are the same for XML files regardless of which connection type is configured in the gem (for example, S3, SFTP, or SharePoint). If you need details on configuring a Source or Target gem end to end (including all tabs such as **Location**), see the documentation for the specific file storage connection. You can also use the [upload file](/data-analysis/gems/source-target/table/upload-files) feature to use XML files. These will be stored in the SQL warehouse configured in your fabric. ## Properties ### Source properties The following properties are available for the XML Source gem. | Property | Description | Default | | ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------- | | Schema | Define the structure of your data. This includes the column names and the data type for each column (such as String, Integer, etc.).

Click **Infer Schema** to automatically detect the structure from the file. | None | | Description | Add a description of the table.

Click **Auto-description** to automatically generate the description. | None | | Row Tag | Specify the XML tag that identifies a single row or record in the dataset. | None | | Inference Data Sampling Limit | Define the maximum number of rows to sample for inferring the schema. Set to `0` to use all rows. | `0` | ### Target properties The following properties are available for the XML Target gem. | Property | Description | Default | | ----------- | ---------------------------------------------------------------------------------------------------------------- | ------- | | Schema | Review the structure of your data. Prophecy automatically detects the schema from the input gem. | None | | Description | Add a description of the table.

Click **Auto-description** to automatically generate the description. | None | ## Schema validation Prophecy lets you enable schema validation for XML files in a Source gem. Use schema validation to ensure XML files conform to a predefined structure before ingesting the data. XML schema validation works for all file storage connections. ### Prerequisites To use schema validation, you need an XSD schema file in the same directory as your source XML file. If your schema uses `xs:include` or `xs:import` to reference external schemas, those referenced schema files must also be in the same directory. ### Set up schema validation 1. In the Location tab, toggle **Enable XSD Schema Validation**. 2. Provide a path to the XSD file that you will use to validate against. 3. Open the Properties tab. 4. Click **Infer Schema**. If the source file schema matches the validation schema, schema inference will run successfully. ### Example The following example shows an XML file and a corresponding XSD schema file that would validate successfully. ```xml users.xml theme={null} 1 John Doe john.doe@example.com ``` ```xml users-schema.xsd theme={null} ``` ### Troubleshooting If the schemas do not match: * Schema inference will fail. * The Source gem will fail to run. To troubleshoot, look for the error in the [runtime logs](/data-analysis/development/runs/runtime-logs). Here is an example error: ```text wrap theme={null} Failed due to error in "reports_xml". Error: XML validation failed against XSD schema for file /path/to/your/file.xml: /tmp/xml_val_1234567890.xml:117: element invalidElement: Schemas validity error : Element '{http://example.com/namespace/v1.0}invalidElement': This element is not expected. Expected is one of ( {http://example.com/namespace/v1.0}validElement1, {http://example.com/namespace/v1.0}validElement2, {http://example.com/namespace/v1.0}validElement3 ). ``` # Google Cloud Storage (GCS) Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/gcs Use Google Cloud Storage (GCS) as a file source or target in a gem This gem runs in . ## Overview Use a Source and Target gem to read from or write to Google Cloud Storage (GCS) locations in Prophecy pipelines. This page covers supported file formats, how to create the gem, and how to configure connection details and paths for both Source and Target gems. ## Supported file formats | Format | Read | Write | | ---------------------------------------------------------------------------- | ---- | ----- | | [CSV](/data-analysis/gems/source-target/file/file-types/csv) | βœ” | βœ” | | [Fixed width](/data-analysis/gems/source-target/file/file-types/fixed-width) | βœ” | | | [JSON](/data-analysis/gems/source-target/file/file-types/json) | βœ” | βœ” | | [Parquet](/data-analysis/gems/source-target/file/file-types/parquet) | βœ” | βœ” | | [XLSX](/data-analysis/gems/source-target/file/file-types/excel) | βœ” | βœ” | | [XML](/data-analysis/gems/source-target/file/file-types/xml) | βœ” | βœ” | ## Create a GCS gem To create a GCS Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **GCS**. Then, click **Next**. In the **Location** tab, set your file format and connection details. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the file properties. These vary based on the file type that you are working with. See the list of properties per file type, such as [CSV](/data-analysis/gems/source-target/file/file-types/csv). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source location When setting up a GCS Source gem, you need to decide how Prophecy should locate files at runtime. There are two modes: * **Filepath**: Always read from a specific file path. You can use wildcards in your path definition. If multiple files match, they are unioned into a single output table. * **Configuration**: Dynamically read files provided by a [file arrival/change trigger](/data-analysis/production/scheduling/triggers) in the pipeline's schedule. When new or updated files are detected in the monitored directory, the trigger starts the pipeline and passes those files to the GCS Source gem, which unions them into a single output table. Use the table below to understand how to configure each option. | Parameter | Description | | ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Select or create connection | Select an existing GCS connection or [create a new one](/data-analysis/environment/connections/gcs). | | Choose Path or Configuration | Choose between the following options.
  • **Filepath**: Read one file from a specified file path.
  • **Configuration**: Dynamically read files provided by a [file arrival/change trigger](/data-analysis/production/scheduling/triggers).
| | File Path
*Filepath option only* | Path to the file in the GCS bucket. Supports wildcards.
Example: `gs://my-bucket/temp/dir/*.csv` | | Examine | Automatically determine the file format, compression type (if any), properties, and schema of the file to read. | | Select Configuration
*Configuration option only* | File arrival/change trigger [configuration](/data-analysis/production/scheduling/triggers#trigger-configuration) that provides the added or modified files for that run. | | Include filename Column | Appends a column containing the source filename for each row in the output table. | | Delete files after successfully processed | Deletes objects after they are successfully read. | | Move files after successfully processed | Moves objects to a specified directory after they are successfully read.

Example: `gs://my-bucket/archive/` | | Format type | Type of file to read, such as `csv` or `json`. | | Compression | The compression type of the file to read.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | ### Configuration If you select **Configuration**, the gem will only run successfully during a **triggered pipeline run**. This is because the gem expects files from the trigger. In other cases, such as an interactive run or an API-triggered run, there will be no files to read. In these situations, you will encounter the following error: ``` Failed due to: Unable to detect modified files for provided File Trigger ``` For the same reason, you'll see an error if you try to infer the schema in the **Properties** tab or load a preview in the **Preview** tab of the Source gem. ## Target location When setting up a GCS Target gem, you need to set the location and file type to correctly write the file. | Parameter | Description | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Select an existing GCS connection or [create a new one](/data-analysis/environment/connections/gcs). | | File Path | GCS path where the output file will be written.
Example: `gs://my-bucket/data/orders.csv` | | Examine | If an existing file exists at the target path, click **Examine** to automatically determine the file format and compression type (if any). | | Encryption Algorithm | The encryption algorithm to use when writing the file.

Supported algorithms: `AES-192`, `AES-296`, `BlowFish` | | Format type | Type of file to write, such as `csv` or `json`. | | Compression | The compression type to use when writing the file.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | Configure encryption in a Target gem to encrypt an entire file. Use the [DataEncoderDecoder](/data-analysis/gems/transform/encoder-decoder) gem to encrypt individual columns. # Microsoft OneDrive file gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/onedrive Use OneDrive as a file source or target in a gem This gem runs in . ## Overview Use a Source or Target gem to read from or write to OneDrive in Prophecy pipelines. This page covers supported file formats, how to create the gem, and how to configure connection details and paths for both Source and Target gems. ## Supported file formats | Format | Read | Write | | ---------------------------------------------------------------------------- | ---- | ----- | | [CSV](/data-analysis/gems/source-target/file/file-types/csv) | βœ” | βœ” | | [Fixed width](/data-analysis/gems/source-target/file/file-types/fixed-width) | βœ” | | | [JSON](/data-analysis/gems/source-target/file/file-types/json) | βœ” | βœ” | | [XLSX](/data-analysis/gems/source-target/file/file-types/excel) | βœ” | βœ” | | [XML](/data-analysis/gems/source-target/file/file-types/xml) | βœ” | βœ” | ## Create a OneDrive gem To create a OneDrive Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **OneDrive**. Then, click **Next**. In the **Location** tab, set your file format and connection details. To learn more, jump to [Location](#location). In the **Properties** tab, set the file properties. These vary based on the file type that you are working with. See the list of properties per file type, such as [CSV](/data-analysis/gems/source-target/file/file-types/csv). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source location | Parameter | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Select an existing OneDrive connection or [create a new one](/data-analysis/environment/connections/onedrive). | | File Path | Path to the file in OneDrive. | | Examine | Automatically determine the file format, compression type (if any), properties, and schema of the file to read. | | Format type | Type of file to read, such as `csv` or `json`. | | Compression | The compression type of the file to read.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | ## Target location | Parameter | Description | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Select an existing OneDrive connection or [create a new one](/data-analysis/environment/connections/onedrive). | | File Path | Path to the file in OneDrive. | | Examine | If an existing file exists at the target path, click **Examine** to automatically determine the file format and compression type (if any). | | Encryption Algorithm | The encryption algorithm to use when writing the file.

Supported algorithms: `AES-192`, `AES-296`, `BlowFish` | | Format type | Type of file to write, such as `csv` or `json`. | | Compression | The compression type to use when writing the file.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | Configure encryption in a Target gem to encrypt an entire file. Use the [DataEncoderDecoder](/data-analysis/gems/transform/encoder-decoder) gem to encrypt individual columns. # Amazon S3 file gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/s3 Use S3 as a file source or target in a gem This gem runs in . ## Overview Use a Source and Target gem to read from or write to S3 locations in Prophecy pipelines. This page covers supported file formats, how to create the gem, and how to configure connection details and paths for both Source and Target gems. ## Supported file formats | Format | Read | Write | | ---------------------------------------------------------------------------- | ---- | ----- | | [CSV](/data-analysis/gems/source-target/file/file-types/csv) | βœ” | βœ” | | [Fixed width](/data-analysis/gems/source-target/file/file-types/fixed-width) | βœ” | | | [JSON](/data-analysis/gems/source-target/file/file-types/json) | βœ” | βœ” | | [Parquet](/data-analysis/gems/source-target/file/file-types/parquet) | βœ” | βœ” | | [XLSX](/data-analysis/gems/source-target/file/file-types/excel) | βœ” | βœ” | | [XML](/data-analysis/gems/source-target/file/file-types/xml) | βœ” | βœ” | ## Create an S3 gem To create an S3 Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **S3**. Then, click **Next**. In the **Location** tab, set your file format and connection details. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the file properties. These vary based on the file type that you are working with. See the list of properties per file type, such as [CSV](/data-analysis/gems/source-target/file/file-types/csv). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source location When setting up an S3 Source gem, you need to decide how Prophecy should locate files at runtime. There are two modes: * **Filepath**: Always read from a specific file path. You can use wildcards in your path definition. If multiple files match, they are unioned into a single output table. * **Configuration**: Dynamically read files provided by a [file arrival/change trigger](/data-analysis/production/scheduling/triggers) in the pipeline's schedule. When new or updated files are detected in the monitored directory, the trigger starts the pipeline and passes those files to the S3 Source gem, which unions them into a single output table. Use the table below to understand how to configure each option. | Parameter | Description | | ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Select or create connection | Select an existing S3 connection or [create a new one](/data-analysis/environment/connections/s3). | | Choose Path or Configuration | Choose between the following options.
  • **Filepath**: Read one file from a specified file path.
  • **Configuration**: Dynamically read files provided by a [file arrival/change trigger](/data-analysis/production/scheduling/triggers).
| | File Path
*Filepath option only* | Path to the file in the S3 bucket. Supports wildcards.
Example: `/temp/dir/*.csv` | | Examine | Automatically determine the file format, compression type (if any), properties, and schema of the file to read. | | Select Configuration
*Configuration option only* | File arrival/change trigger [configuration](/data-analysis/production/scheduling/triggers#trigger-configuration) that provides the added or modified files for that run. | | Include filename Column | Appends a column containing the source filename for each row in the output table. | | Delete files after successfully processed | Deletes objects after they are successfully read. | | Move files after successfully processed | Moves objects to a specified directory after they are successfully read. | | Format type | Type of file to read, such as `csv` or `json`. | | Compression | The compression type of the file to read.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | ### Configuration If you select **Configuration**, the gem will only run successfully during a **triggered pipeline run**. This is because the gem expects files from the trigger. In other cases, such as an interactive run or an API-triggered run, there will be no files to read. In these situations, you will encounter the following error: ``` Failed due to: Unable to detect modified files for provided File Trigger ``` For the same reason, you'll see an error if you try to infer the schema in the **Properties** tab or load a preview in the **Preview** tab of the Source gem. ## Target location When setting up an S3 Target gem, you need to set the location and file type to correctly write the file. | Parameter | Description | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Select an existing S3 connection or [create a new one](/data-analysis/environment/connections/s3). | | File Path | S3 path where the output file will be written.
Example: `s3://my-bucket/data/orders.csv` | | Examine | If an existing file exists at the target path, click **Examine** to automatically determine the file format and compression type (if any). | | Encryption Algorithm | The encryption algorithm to use when writing the file.

Supported algorithms: `AES-192`, `AES-296`, `BlowFish` | | Format type | Type of file to write, such as `csv` or `json`. | | Compression | The compression type to use when writing the file.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | Configure encryption in a Target gem to encrypt an entire file. Use the [DataEncoderDecoder](/data-analysis/gems/transform/encoder-decoder) gem to encrypt individual columns. # SFTP file gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/sftp Use SFTP as a file source or target in a gem This gem runs in . ## Overview Use a Source and Target gem to read from or write to remote SFTP locations in Prophecy pipelines. This page covers supported file formats, how to create the gem, and how to configure connection details and paths for both Source and Target gems. ## Supported file formats The SFTP gem supports the following file formats. | Format | Read | Write | | ---------------------------------------------------------------------------- | ---- | ----- | | [CSV](/data-analysis/gems/source-target/file/file-types/csv) | βœ” | βœ” | | [Fixed width](/data-analysis/gems/source-target/file/file-types/fixed-width) | βœ” | | | [JSON](/data-analysis/gems/source-target/file/file-types/json) | βœ” | βœ” | | [Parquet](/data-analysis/gems/source-target/file/file-types/parquet) | βœ” | βœ” | | [XML](/data-analysis/gems/source-target/file/file-types/xml) | βœ” | βœ” | | [XLSX](/data-analysis/gems/source-target/file/file-types/excel) | βœ” | βœ” | ## Create an SFTP gem To create an SFTP Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **SFTP**. Then, click **Next**. In the **Location** tab, set your file format and connection details. To learn more, jump to [Source location](#source-location) and [Target location](#target-location). In the **Properties** tab, set the file properties. These vary based on the file type that you are working with. See the list of properties per file type, such as [CSV](/data-analysis/gems/source-target/file/file-types/csv). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source location When setting up an SFTP Source gem, you need to decide how Prophecy should locate files at runtime. There are two modes: * **Filepath**: Always read from a specific file path. You can use wildcards in your path definition. If multiple files match, they are unioned into a single output table. * **Configuration**: Dynamically read files provided by a [file arrival/change trigger](/data-analysis/production/scheduling/triggers) in the pipeline's schedule. When new or updated files are detected in the monitored directory, the trigger starts the pipeline and passes those files to the SFTP Source gem, which unions them into a single output table. Use the table below to understand how to configure each option. | Parameter | Description | | ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Select or create connection | Choose an existing SFTP connection or [create a new one](/data-analysis/environment/connections/sftp). | | Choose Path or Configuration | Choose between the following options.
  • **Filepath**: Read one file from a specified file path.
  • **Configuration**: Dynamically read files provided by a [file arrival/change trigger](/data-analysis/production/scheduling/triggers).
| | File Path
*Filepath option only* | Path to the file on the SFTP server. Supports wildcards.
Example: `/temp/dir/*.csv` | | Examine | Automatically determine the file format, compression type (if any), properties, and schema of the file to read. | | Select Configuration
*Configuration option only* | File arrival/change trigger [configuration](/data-analysis/production/scheduling/triggers#trigger-configuration) that provides the added or modified files for that run. | | Include filename Column | Appends a column containing the source filename for each row in the output table. | | Delete files after successfully processed | Deletes files after they are successfully read. | | Move files after successfully processed | Moves files to a specified directory after they are successfully read. | | Format type | Type of file to read, such as `csv` or `json`. | | Compression | The compression type of the file to read.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | ### Configuration If you select **Configuration**, the gem will only run successfully during a **triggered pipeline run**. This is because the gem expects files from the trigger. In other cases, such as an interactive run or an API-triggered run, there will be no files to read. In these situations, you might encounter the following error: ``` Failed due to: Unable to detect modified files for provided File Trigger ``` For the same reason, you'll see an error if you try to infer the schema in the **Properties** tab or load a preview in the **Preview** tab of the Source gem. ## Target location When setting up an SFTP Target gem, you need to set the location and file type to correctly write the file. | Parameter | Description | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Select an existing SFTP connection or [create a new one](/data-analysis/environment/connections/sftp). | | File Path | SFTP path where the output file will be written.
Example: `/temp/dir/file.csv` | | Examine | If an existing file exists at the target path, click **Examine** to automatically determine the file format and compression type (if any). | | Encryption Algorithm | The encryption algorithm to use when writing the file.

Supported algorithms: `AES-192`, `AES-296`, `BlowFish` | | Format type | Type of file to write, such as `csv` or `json`. | | Compression | The compression type to use when writing the file.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | Configure encryption in a Target gem to encrypt an entire file. Use the [DataEncoderDecoder](/data-analysis/gems/transform/encoder-decoder) gem to encrypt individual columns. # Microsoft SharePoint file gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/sharepoint Use SharePoint as a file source or target in a gem This gem runs in . ## Overview Use a Source or Target gem to read from or write to SharePoint in Prophecy pipelines. This page covers supported file formats, how to create the gem, and how to configure connection details and paths for both Source and Target gems. ## Supported file formats | Format | Read | Write | | ---------------------------------------------------------------------------- | ---- | ----- | | [CSV](/data-analysis/gems/source-target/file/file-types/csv) | βœ” | βœ” | | [Fixed width](/data-analysis/gems/source-target/file/file-types/fixed-width) | βœ” | | | [JSON](/data-analysis/gems/source-target/file/file-types/json) | βœ” | βœ” | | [XLSX](/data-analysis/gems/source-target/file/file-types/excel) | βœ” | βœ” | | [XML](/data-analysis/gems/source-target/file/file-types/xml) | βœ” | βœ” | ## Add a SharePoint Source or Target gem To add a SharePoint Source or Target gem to your pipeline: 1. Open your pipeline in [Studio](/data-analysis/development/studio/studio). 2. Click **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click the gem to configure it. In the **Type** tab, select **SharePoint**. Click **Next**. In the **Location** tab, set your file format and connection details. To learn more, see [Location](#location) below. In the **Properties** tab, set file properties. These vary based on the file type that you are working with. See the list of properties per file type, such as [CSV](/data-analysis/gems/source-target/file/file-types/csv). In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Grant service principal access to SharePoint sites The connection you select or create in the **Location** tab authenticates to SharePoint as a service principal (an Azure AD app registration). Before that connection can read from or write to a site, the service principal needs explicit access to it. Microsoft recommends the `Sites.Selected` permission over tenant-wide access, since it scopes the service principal to specific site collections only. Adding `Sites.Selected` in Azure alone does not grant access to any site. You must also complete [Step 2](#step-2-get-the-sharepoint-site-id) and [Step 3](#step-3-grant-access-to-the-site) below for each site the connection needs. ### Prerequisites * An Azure AD app registration to use as the service principal. * SharePoint Administrator or Global Administrator role on the tenant. * Access to [Microsoft Graph Explorer](https://developer.microsoft.com/en-us/graph/graph-explorer). 1. In the [Azure Portal](https://portal.azure.com), go to **App Registrations** and select your app. 2. Click **API permissions** in the left menu. 3. Click **Add a permission** > **Microsoft Graph** > **Application permissions**. 4. Search for and add `Sites.Selected`. 5. Click **Grant admin consent for your tenant**. | Permission | Type | Description | | ---------------- | ----------- | -------------------------------- | | `Sites.Selected` | Application | Access selected site collections | Use Graph Explorer to look up the ID of the site you want to grant access to. 1. Go to [Graph Explorer](https://developer.microsoft.com/en-us/graph/graph-explorer) and sign in with a SharePoint Administrator or Global Administrator account. 2. Take the site's normal URL (for example, `https://prophecy.sharepoint.com/sites/testsite`) and convert it to the Graph format: replace the first `/` after the hostname with a colon. 3. Set the method to `GET` and run: ``` GET https://graph.microsoft.com/v1.0/sites/{hostname}:/sites/{siteName} ``` For example: ``` GET https://graph.microsoft.com/v1.0/sites/prophecy.sharepoint.com:/sites/testsite ``` 4. Copy the `id` value from the response. This is the site ID used in the next step. The hostname and site path are joined with a colon, not a slash. * Wrong: `prophecy.sharepoint.com/sites/testsite` * Correct: `prophecy.sharepoint.com:/sites/testsite` Use a `POST` request to assign the service principal a role on the specific site. 1. In Graph Explorer, set the method to `POST` and use the site ID from Step 2: ``` POST https://graph.microsoft.com/v1.0/sites/{site-id}/permissions ``` 2. In the **Request body** tab, paste the following, replacing `` and `` with your service principal's values: ```json theme={null} { "roles": ["read", "write"], "grantedToIdentities": [ { "application": { "id": "", "displayName": "" } } ] } ``` 3. Click **Run query**. A `201 Created` response confirms the permission was granted. Repeat this step for every additional site the connection needs access to. | Role value | Access level | | ------------- | ----------------- | | `read` | Read-only | | `write` | Read and write | | `fullcontrol` | Full site control | Grant both `read` and `write` together, as shown above. This is the only combination Prophecy has tested; granting `read` alone for Source-only connections is untested and not currently recommended. `Sites.Selected` grants access at the site level only. It does not support scoping access to individual folders or libraries within a site. ### Troubleshooting #### 403 Forbidden when running the POST request Confirm you're signed in to Graph Explorer as a SharePoint Administrator or Global Administrator. In Graph Explorer, open **Modify permissions** and consent to `Sites.FullControl.All`, then re-run the request. #### Invalid hostname error Make sure that you have a colon `:` before `/sites` in the hostname. Correct: `prophecy.sharepoint.com:/sites/testsite` (uses colon before /sites/) Wrong: `prophecy.sharepoint.com/sites/testsite` (uses slash) #### Permissions don't seem to take effect\*\* * `Sites.Selected` grants nothing by default; you need to grant access per site explicitly via Step 3. * Confirm Step 3 was repeated for each site the connection needs, not just the first one. * Permission grants can take a few minutes to propagate. If a connection test fails immediately after granting access, wait a few minutes and retry before troubleshooting further. **Check what's already been granted** Run a `GET` on the site's permissions to see existing grants, including any stale or duplicate entries from a previous app registration: ``` GET https://graph.microsoft.com/v1.0/sites/{site-id}/permissions ``` ## Source location | Parameter | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Select or create connection | Select an existing SharePoint connection or [create a new one](/data-analysis/environment/connections/sharepoint). See [Grant service principal access to SharePoint sites](#grant-service-principal-access-to-sharepoint-sites) if the connection needs new site permissions. | | File Path | Path to the file in SharePoint.

Example: `/sites/sitename/SharedDocuments/file.csv` | | Examine | Automatically determine the file format, compression type (if any), properties, and schema of the file to read. | | Format type | Type of file to read, such as `csv` or `json`. | | Compression | The compression type of the file to read.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | ## Target location | Parameter | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Select or create connection | Select an existing SharePoint connection or [create a new one](/data-analysis/environment/connections/sharepoint). See [Grant service principal access to SharePoint sites](#grant-service-principal-access-to-sharepoint-sites) if the connection needs new site permissions. | | File Path | Path to the file in SharePoint.

Example: `/sites/sitename/SharedDocuments/file.csv` | | Examine | If an existing file exists at the target path, click **Examine** to automatically determine the file format and compression type (if any). | | Encryption Algorithm | The encryption algorithm to use when writing the file.

Supported algorithms: `AES-192`, `AES-296`, `BlowFish` | | Format type | Type of file to write, such as `csv` or `json`. | | Compression | The compression type to use when writing the file.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | Configure encryption in a Target gem to encrypt an entire file. Use the [DataEncoderDecoder](/data-analysis/gems/transform/encoder-decoder) gem to encrypt individual columns. # Smartsheet file gem Source: https://docs.prophecy.ai/data-analysis/gems/source-target/file/smartsheet Use Smartsheet as a file source or target in a gem This gem runs in . ## Overview Use a Source and Target gem to read from or write to Smartsheet locations in Prophecy pipelines. This page covers supported file formats, how to create the gem, and how to configure connection details and paths for both Source and Target gems. ## Supported file formats | Format | Read | Write | | --------------------------------------------------------------- | ---- | ----- | | [XLSX](/data-analysis/gems/source-target/file/file-types/excel) | βœ” | βœ” | ## Create a Smartsheet gem To create a Smartsheet Source or Target gem in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Source** or **Target** from the dropdown. 4. Click on the gem to open the configuration. In the **Type** tab, select **Smartsheet**. Then, click **Next**. In the **Location** tab, set your file format and connection details. To learn more, jump to [Location](#location). In the **Properties** tab, set the file properties. See [Excel](/data-analysis/gems/source-target/file/file-types/excel) to view XLSX-specific properties. In the **Preview** tab, load a sample of the data and verify that it looks correct. ## Source location | Parameter | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Select an existing Smartsheet connection or [create a new one](/data-analysis/environment/connections/smartsheet). | | File Path | Path to the file in Smartsheet.
Example: `/Projects/ProjectName/file.csv` | | Examine | Click **Examine** to automatically determine the file format and compression type (if any). | | Format type | Type of file to read, such as `csv` or `json`. | | Compression | The compression type of the file to read.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | ## Target location | Parameter | Description | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Select or create connection | Select an existing Smartsheet connection or [create a new one](/data-analysis/environment/connections/smartsheet). | | File Path | Path to the file in Smartsheet.
Example: `/Projects/ProjectName/file.csv` | | Examine | If an existing file exists at the target path, click **Examine** to automatically determine the file format and compression type (if any). | | Encryption Algorithm | The encryption algorithm to use when writing the file.

Supported algorithms: `AES-192`, `AES-296`, `BlowFish` | | Format type | Type of file to write, such as `csv` or `json`. | | Compression | The compression type to use when writing the file.

Supported types: `uncompressed`,`gzip`, `zstd`, `lz4`, `zlib`, `snappy`, `lzop` | Configure encryption in a Target gem to encrypt an entire file. Use the [DataEncoderDecoder](/data-analysis/gems/transform/encoder-decoder) gem to encrypt individual columns. # Source and target types Source: https://docs.prophecy.ai/data-analysis/gems/source-target/source-target Read data into and write data out of your pipelines Table, Source, and Target gems define how Prophecy pipelines read data from a source or writing data to a destination. These gems handle interactions with both external systems and the [SQL Warehouse Connection](/data-analysis/environment/fabrics/prophecy-fabrics) defined in your Prophecy fabric. Regardless of gem type, none of the data you read and write in a pipeline is persisted in Prophecy. All data is transformed in memory, and no data gets written to disk. ## Gem types Prophecy supports different gem types for reading from and writing to data sources. | Gem type | Data provider | Description | | ---------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Table gem | Data warehouse | Represents datasets in the in the [SQL Warehouse Connection](/data-analysis/environment/fabrics/prophecy-fabrics) of a Prophecy fabric. Tables can act as sources, targets, or intermediate stages in a pipeline. | | Source gem | External system | Represents data (tables or files) stored in external platforms outside the SQL warehouse. Source gems provide input to a pipeline. | | Target gem | External system | Represents data (tables or files) stored in external platforms outside the SQL warehouse. Target gems consume output from a pipeline. | Source and target gems cannot serve as intermediate steps in a pipeline. Source gems only have output ports, and Target gems only have input ports. Pipeline performance depends on how your connections are configured. Processing native tables within the warehouse is more efficient than accessing external sources. Avoid setting up an external connection that points to the same warehouse defined in your fabric. This redundancy can lead to slower performance. ## Data formats Prophecy supports the following data formats: * Tables from the connected SQL warehouse. These are accessed with Table gems. * Tables from external systems. These are accessed with Source and Target gems. * Files from external systems. These are accessed with Source and Target gems. To use external systems, you need to set up corresponding [connections](/data-analysis/environment/connections/connections). # Read from BigQuery tables Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/bigquery-read Configure a BigQuery table, view, or seed as a read source This gem runs in . ## Overview In Prophecy, datasets stored in the [SQL Warehouse Connection](/data-analysis/environment/fabrics/prophecy-fabrics) defined in your fabric are accessed using Table gems. Unlike Source and Target gems, Table gems run directly within the data warehouse, eliminating extra orchestration steps and improving performance. This page explains how to read from a BigQuery table, view, or seed using the Table gem. To write to a BigQuery table or view instead, see [Write to BigQuery tables](/data-analysis/gems/source-target/table/bigquery-write). ## Table types The following table types are supported for BigQuery connections. | Name | Description | Type | | ----- | ------------------------------------------------------------------------------------------------------------- | ---------------- | | Table | Persistent storage of structured data in your SQL warehouse. Faster for frequent queries (indexed). | Source or Target | | View | A virtual table that derives data dynamically from a query. Slower for complex queries (computed at runtime). | Source or Target | | Seed | Small CSV-format files that you can write directly in Prophecy. | **Source only** | ## Configure table Once you create a Table gem, you can reuse it throughout your project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To read from a table in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Table** from the dropdown. 4. Click on the gem to open the configuration. Choose the table, view, or seed you want to read from the list. Confirm the type: **Table**, **View**, or **Seed**. Seeds configure differently from Tables and Views β€” they skip the Location step entirely. See [Configure seeds](#configure-seeds) below. The Location tab defines where a table lives and how it is identified within your project. | Field | Description | | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Table alias | A stable, logical identifier for the table that stays constant even if the underlying database, schema, or table name changes. Required when you create the table, and cannot be changed afterward. | | Database | The database containing the table. | | Schema | The schema containing the table. | | Table | The table name. | A **parameter set** selector appears in the upper-right corner of the Location tab: * If the table gem is used inside a pipeline, the selector shows that pipeline's active parameter set automatically. * If you're working on the table outside of a pipeline (for example, from the Project browser), the selector shows **Select Pipeline and Parameter set** until you choose one. You only need to do this if one or more Location fields are set to Advanced mode (see below) and you need their values resolved outside pipeline context. At the bottom of the Location tab, Prophecy shows a live preview of the resolved `database.schema.table` location: the hardcoded values if all fields are in Simple mode, or the resolved parameter values if any field is in Advanced mode. If a value can't be resolved yet β€” for example, because no parameter set is selected β€” Prophecy displays the raw value instead. #### Make a location field dynamic In the default **Simple** mode, each Location field (database, schema, table) takes a fixed value that you type directly. Switch a field to **Advanced** mode to bind it to a project or pipeline parameter instead of a fixed value. Only parameters of type `sql_expression` can be used in Advanced mode. Using a parameter of a different type will cause the table location to fail to resolve. Switching a field from Advanced back to Simple mode clears its current value. Once a field is in Advanced mode, its value depends on which parameter set is active (see the parameter set selector above). This makes it possible to define a table once and reuse it across multiple pipelines, each supplying different values for the parameterized fields via their own parameter sets. To reuse a table you've already created, select it from **Table > \[alias]** in the Project browser. Prophecy does not validate that tables resolved from different parameter sets share the same schema. If your parameter sets point to tables with different schemas, downstream steps in your pipeline may fail or behave unexpectedly. Define or infer the schema. Add a description if needed. If your BigQuery tables are partitioned or clustered, Prophecy automatically displays the partitioning and clustering information in the Properties tab. You'll be able to view information such as the partition granularity, partitioning data type, clustering columns, and more. You won't be able to edit these properties. #### Working with JSON columns Prophecy supports BigQuery `JSON` columns, including schema inference, nested field exploration, and reading and writing JSON data. BigQuery JSON columns support schema inference and nested field access similar to Snowflake `VARIANT` columns. ##### Infer schema For tables that contain `JSON` columns, use **Infer Schema** to sample data and discover the nested structure stored within the column. The inferred schema is displayed as an expandable tree in the Properties tab, similar to Snowflake `VARIANT` columns. For example, a JSON column containing: ```json theme={null} { "amount": 100, "user": { "id": 123, "plan": "premium" } } ``` is displayed as nested fields that can be expanded and referenced throughout your pipeline. ##### Access nested fields After schema inference, you can reference nested JSON fields in gems such as Reformat, Filter, and Join using dot notation. For example: ```text theme={null} payload.user.plan ``` Prophecy automatically generates the appropriate BigQuery JSON functions: | Field type | Generated SQL | | ----------------- | -------------------------------------------------- | | String values | `JSON_VALUE(payload, '$.user.plan')` | | Numeric values | `CAST(JSON_VALUE(payload, '$.amount') AS FLOAT64)` | | Boolean values | `CAST(JSON_VALUE(payload, '$.active') AS BOOL)` | | Objects or arrays | `JSON_QUERY(payload, '$.user')` | Writing to a JSON column has different behavior β€” Prophecy auto-converts string JSON to native BigQuery JSON values on write. See [Write to BigQuery tables](/data-analysis/gems/source-target/table/bigquery-write) for details. Load a sample of the data before saving. For views, this loads data based on the view's underlying query. Add data tests to validate the data you're reading. See [Table tests vs. project tests](/data-analysis/development/tests/test-comparison) to decide which approach fits your validation needs. ## Configure seeds Seeds are lightweight CSV datasets defined in your project. Seeds are source-only and don't support writing. | Parameter | Description | | ---------- | ---------------------------------------------------------------------------------------------------------------------- | | Properties | Copy-paste your CSV data and define certain [properties](https://docs.getdbt.com/reference/seed-configs) of the table. | | Preview | Load a preview of your seed in table format. | Seeds are implemented as [dbt seeds](https://docs.getdbt.com/docs/build/seeds) under the hood. The CSV data you define is stored in your Prophecy project files and materialized as a table in your data warehouse. This table is created in the [default target dataset](/data-analysis/environment/connections/bigquery#connection-parameters) specified in your BigQuery connection. ## Reusing and sharing tables After you create a table in Prophecy, you can reuse its configuration across your entire project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To make tables available to other teams, you can share your project as a package in the [Package Hub](/data-analysis/development/extensibility/package-hub/package-hub). Other users will be able to use the shared table configuration, provided they have the necessary permissions in BigQuery to access the underlying data. # Write to BigQuery tables Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/bigquery-write Configure a BigQuery table or view as a write target This gem runs in . ## Overview In Prophecy, datasets stored in the [SQL Warehouse Connection](/data-analysis/environment/fabrics/prophecy-fabrics) defined in your fabric are accessed using Table gems. Unlike Source and Target gems, Table gems run directly within the data warehouse, eliminating extra orchestration steps and improving performance. This page explains how to write to a BigQuery table or view using the Table gem. To read from an existing BigQuery table or view instead, see [Read from BigQuery tables](/data-analysis/gems/source-target/table/bigquery-read). ## Configure table Once you create a Table gem, you can reuse it throughout your project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To write to a table in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Table** from the dropdown. 4. Click on the gem to open the configuration. To write to an existing table, select it from the list. To write to a new table, click **+ New Table**. Choose **Table** or **View**. The Location tab defines where a table lives and how it is identified within your project. | Field | Description | | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Table alias | A stable, logical identifier for the table that stays constant even if the underlying database, schema, or table name changes. Required when you create the table, and cannot be changed afterward. | | Database | The database containing the table. | | Schema | The schema containing the table. | | Table | The table name. | A **parameter set** selector appears in the upper-right corner of the Location tab: * If the table gem is used inside a pipeline, the selector shows that pipeline's active parameter set automatically. * If you're working on the table outside of a pipeline (for example, from the Project browser), the selector shows **Select Pipeline and Parameter set** until you choose one. You only need to do this if one or more Location fields are set to Advanced mode (see below) and you need their values resolved outside pipeline context. At the bottom of the Location tab, Prophecy shows a live preview of the resolved `database.schema.table` location: the hardcoded values if all fields are in Simple mode, or the resolved parameter values if any field is in Advanced mode. If a value can't be resolved yet β€” for example, because no parameter set is selected β€” Prophecy displays the raw value instead. #### Make a location field dynamic In the default **Simple** mode, each Location field (database, schema, table) takes a fixed value that you type directly. Switch a field to **Advanced** mode to bind it to a project or pipeline parameter instead of a fixed value. Only parameters of type `sql_expression` can be used in Advanced mode. Using a parameter of a different type will cause the table location to fail to resolve. Switching a field from Advanced back to Simple mode clears its current value. Once a field is in Advanced mode, its value depends on which parameter set is active (see the parameter set selector above). This makes it possible to define a table once and reuse it across multiple pipelines, each supplying different values for the parameterized fields via their own parameter sets. To reuse a table you've already created, select it from **Table > \[alias]** in the Project browser. Prophecy does not validate that tables resolved from different parameter sets share the same schema. If your parameter sets point to tables with different schemas, downstream steps in your pipeline may fail or behave unexpectedly. Map each incoming column to a column on the target table. Prophecy suggests mappings automatically β€” review and adjust them as needed. Any target column left unmapped defaults to null. This is schema **mapping**, not schema definition or inference. Prophecy already knows the target's schema (from an existing table, or as you define one for a new table in the Location step); this step is about reconciling your pipeline's output columns against it. You can also set a description for the table and configure generic options, such as skipping execution when the input has zero rows. #### Working with JSON columns Prophecy supports BigQuery `JSON` columns, including schema inference, nested field exploration, and reading and writing JSON data. When writing to a BigQuery `JSON` column, Prophecy automatically converts string representations of JSON into native BigQuery JSON values β€” no manual SQL transformation is required. BigQuery JSON columns support schema inference and nested field access similar to Snowflake `VARIANT` columns. ### Map schema for existing target tables When you select an existing table as a target, the incoming schema might not match the schema of the target table. Prophecy lets you reconcile these differences in the **Map Schema** section of the Properties tab. For each target column, select the corresponding source column. Prophecy can suggest mappings for unmapped columns, which you can review and select. Any required casts or other transformations are applied so that the incoming data conforms to the target schema. If you want the target table to use the incoming schema instead, click **Overwrite Target Schema**. This replaces the existing target schema with the source schema rather than mapping the incoming columns to it. Schema mapping is available when writing to existing target tables in Snowflake, Databricks, and BigQuery. Select how you want the data to be written each time you run the pipeline. Learn more in [Write strategies](/data-analysis/gems/source-target/table/write/write-options). This step applies to **Table** targets only. If you selected **View** in the type & format step, there's no Write Options tab β€” the gem shows **Preview** instead, even when it's positioned at the end of a pipeline. Views are always fully recomputed and overwritten on each run, so there's no write mode to choose. Add data tests to validate the table after it's written. See [Table tests vs. project tests](/data-analysis/development/tests/test-comparison) to decide which approach fits your validation needs. ## Reusing and sharing tables After you create a table in Prophecy, you can reuse its configuration across your entire project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To make tables available to other teams, you can share your project as a package in the [Package Hub](/data-analysis/development/extensibility/package-hub/package-hub). Other users will be able to use the shared table configuration, provided they have the necessary permissions in BigQuery to access the underlying data. # Read from Databricks tables Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/databricks-read Configure a Databricks table, view, or seed as a read source This gem runs in . ## Overview In Prophecy, datasets stored in the [SQL Warehouse Connection](/data-analysis/environment/fabrics/prophecy-fabrics) defined in your fabric are accessed using Table gems. Unlike Source and Target gems, Table gems run directly within the data warehouse, eliminating extra orchestration steps and improving performance. This page explains how to read from a Databricks table, view, or seed using the Table gem. To write to a Databricks table or view instead, see [Write to Databricks tables](/data-analysis/gems/source-target/table/databricks-write). ## Table types The following table types are supported for Databricks connections. | Name | Description | Type | | ----- | ------------------------------------------------------------------------------------------------------------- | ---------------- | | Table | Persistent storage of structured data in your SQL warehouse. Faster for frequent queries (indexed). | Source or Target | | View | A virtual table that derives data dynamically from a query. Slower for complex queries (computed at runtime). | Source or Target | | Seed | Small CSV-format files that you can write directly in Prophecy. | **Source only** | For more information, visit the Databricks documentation on [Tables](https://docs.databricks.com/aws/en/tables/table-overview) and [Views](https://docs.databricks.com/aws/en/views/). ## Configure table Once you create a Table gem, you can reuse the table throughout your project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To read from a table in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Table** from the dropdown. 4. Click on the gem to open the configuration. Choose the table, view, or seed you want to read from the list. Seeds configure differently from Tables and Views β€” see [Configure seeds](#configure-seeds) below. The Location tab defines where a table lives and how it is identified within your project. | Field | Description | | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Table alias | A stable, logical identifier for the table that stays constant even if the underlying database, schema, or table name changes. Required when you create the table, and cannot be changed afterward. | | Database | The database containing the table. | | Schema | The schema containing the table. | | Table | The table name. | A **parameter set** selector appears in the upper-right corner of the Location tab: * If the table gem is used inside a pipeline, the selector shows that pipeline's active parameter set automatically. * If you're working on the table outside of a pipeline (for example, from the Project browser), the selector shows **Select Pipeline and Parameter set** until you choose one. You only need to do this if one or more Location fields are set to Advanced mode (see below) and you need their values resolved outside pipeline context. At the bottom of the Location tab, Prophecy shows a live preview of the resolved `database.schema.table` location: the hardcoded values if all fields are in Simple mode, or the resolved parameter values if any field is in Advanced mode. If a value can't be resolved yet β€” for example, because no parameter set is selected β€” Prophecy displays the raw value instead. #### Make a location field dynamic In the default **Simple** mode, each Location field (database, schema, table) takes a fixed value that you type directly. Switch a field to **Advanced** mode to bind it to a project or pipeline parameter instead of a fixed value. Only parameters of type `sql_expression` can be used in Advanced mode. Using a parameter of a different type will cause the table location to fail to resolve. Switching a field from Advanced back to Simple mode clears its current value. Once a field is in Advanced mode, its value depends on which parameter set is active (see the parameter set selector above). This makes it possible to define a table once and reuse it across multiple pipelines, each supplying different values for the parameterized fields via their own parameter sets. To reuse a table you've already created, select it from **Table > \[alias]** in the Project browser. Prophecy does not validate that tables resolved from different parameter sets share the same schema. If your parameter sets point to tables with different schemas, downstream steps in your pipeline may fail or behave unexpectedly. For a **View**, enter the database, schema, and view name here instead of a table name. Click **Infer Schema** to infer the table's schema. We recommend keeping values as generated, because these match column names and types in your SQL warehouse. When you click **Infer Schema**, Prophecy also generates a description for the table. You can edit the description by clicking the field. For a **View**, this defines or infers the schema from the view's underlying query. Click **Preview** to view a sample of the table's data. For a **View**, this loads data based on the view's underlying query rather than stored data. Here, you can create and run [table tests](/data-analysis/development/tests/table-tests) for the table. Table tests are reusable, parameterized SQL queries that validate your data quality. See [Table tests vs. project tests](/data-analysis/development/tests/test-comparison) if you're deciding which approach fits your validation needs. When you are satisfied with the table's configuration, click **Save**. ## Configure seeds Seeds are lightweight CSV datasets defined in your project. Seeds are source-only. | Parameter | Description | | ---------- | ---------------------------------------------------------------------------------------------------------------------- | | Properties | Copy-paste your CSV data and define certain [properties](https://docs.getdbt.com/reference/seed-configs) of the table. | | Preview | Load a preview of your seed in table format. | Seeds are implemented as [dbt seeds](https://docs.getdbt.com/docs/build/seeds) under the hood. The CSV data you define is stored in your Prophecy project files and materialized as a table in your data warehouse. This table is created in the [default target schema](/data-analysis/environment/connections/databricks#connection-parameters) specified in your Databricks connection. Tables in pipelines do not support dbt properties, which are only applicable to [model sources and targets](/data-engineering/development/models/sources-target/sources-target). The **properties** referenced above are the seed's own properties (the dbt seed config), not pipeline table properties. ## Cross-workspace access If your fabric uses Databricks as the SQL warehouse, you can't select Databricks in an external Source or Target gem. Instead, you must use Table gems, which are limited to the Databricks warehouse defined in the SQL warehouse connection. To work with tables from a different Databricks workspace, use [Delta Sharing](https://docs.databricks.com/aws/en/delta-sharing/). Delta Sharing lets you access data across workspaces without creating additional Databricks connections. Prophecy implements this guardrail to avoid using external connections when the data can be made available in your warehouse. External connections introduce an extra data transfer step, which slows down pipeline execution and adds unnecessary complexity. For best performance, Prophecy always prefers reading and writing directly within the warehouse. ## Reusing and sharing tables After you create a table in Prophecy, you can reuse its configuration across your entire project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To make tables available to other teams, you can share your project as a package in the [Package Hub](/data-analysis/development/extensibility/package-hub/package-hub). Other users will be able to use the shared table configuration, provided they have the necessary permissions in Databricks to access the underlying data. # Write to Databricks tables Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/databricks-write Configure a Databricks table or view as a write target This gem runs in . ## Overview In Prophecy, datasets stored in the [SQL Warehouse Connection](/data-analysis/environment/fabrics/prophecy-fabrics) defined in your fabric are accessed using Table gems. Unlike Source and Target gems, Table gems run directly within the data warehouse, eliminating extra orchestration steps and improving performance. This page explains how to write to a Databricks table or view using the Table gem. To read from an existing Databricks table or view instead, see [Read from Databricks tables](/data-analysis/gems/source-target/table/write/write-options). ## Configure table Once you create a Table gem, you can reuse it throughout your project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To write to a table in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Table** from the dropdown. 4. Click on the gem to open the configuration. To write to an existing table, select it from the list. To write to a new table, click **+ New Table**. Choose **Table** or **View**. Seeds are read-only and aren't available as a write target. The Location tab defines where a table lives and how it is identified within your project. | Field | Description | | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Table alias | A stable, logical identifier for the table that stays constant even if the underlying database, schema, or table name changes. Required when you create the table, and cannot be changed afterward. | | Database | The database containing the table. | | Schema | The schema containing the table. | | Table | The table name. | A **parameter set** selector appears in the upper-right corner of the Location tab: * If the table gem is used inside a pipeline, the selector shows that pipeline's active parameter set automatically. * If you're working on the table outside of a pipeline (for example, from the Project browser), the selector shows **Select Pipeline and Parameter set** until you choose one. You only need to do this if one or more Location fields are set to Advanced mode (see below) and you need their values resolved outside pipeline context. At the bottom of the Location tab, Prophecy shows a live preview of the resolved `database.schema.table` location: the hardcoded values if all fields are in Simple mode, or the resolved parameter values if any field is in Advanced mode. If a value can't be resolved yet β€” for example, because no parameter set is selected β€” Prophecy displays the raw value instead. #### Make a location field dynamic In the default **Simple** mode, each Location field (database, schema, table) takes a fixed value that you type directly. Switch a field to **Advanced** mode to bind it to a project or pipeline parameter instead of a fixed value. Only parameters of type `sql_expression` can be used in Advanced mode. Using a parameter of a different type will cause the table location to fail to resolve. Switching a field from Advanced back to Simple mode clears its current value. Once a field is in Advanced mode, its value depends on which parameter set is active (see the parameter set selector above). This makes it possible to define a table once and reuse it across multiple pipelines, each supplying different values for the parameterized fields via their own parameter sets. To reuse a table you've already created, select it from **Table > \[alias]** in the Project browser. Prophecy does not validate that tables resolved from different parameter sets share the same schema. If your parameter sets point to tables with different schemas, downstream steps in your pipeline may fail or behave unexpectedly. Map each incoming column to a column on the target table. Prophecy suggests mappings automatically β€” review and adjust them as needed. Any target column left unmapped defaults to null. This is schema **mapping**, not schema definition or inference. Prophecy already knows the target's schema (from an existing table, or as you define one for a new table in the Location step); this step is about reconciling your pipeline's output columns against it. You can also set a description for the table and configure generic options, such as skipping execution when the input has zero rows. ### Map schema for existing target tables When you select an existing table as a target, the incoming schema might not match the schema of the target table. Prophecy lets you reconcile these differences in the **Map Schema** section of the Properties tab. For each target column, select the corresponding source column. Prophecy can suggest mappings for unmapped columns, which you can review and select. Any required casts or other transformations are applied so that the incoming data conforms to the target schema. If you want the target table to use the incoming schema instead, click **Overwrite Target Schema**. This replaces the existing target schema with the source schema rather than mapping the incoming columns to it. Schema mapping is available when writing to existing target tables in Snowflake, Databricks, and BigQuery. Select how you want the data to be written each time you run the pipeline. Learn more in [Write strategies](/data-analysis/gems/source-target/table/write/write-options). This step applies to **Table** targets only. If you selected **View** in the type & format step, there's no Write Options tab β€” the gem shows **Preview** instead, even when it's positioned at the end of a pipeline. Views are always fully recomputed and overwritten on each run, so there's no write mode to choose. Add data tests to validate the table after it's written. See [Table tests vs. project tests](/data-analysis/development/tests/test-comparison) to decide which approach fits your validation needs. ## Cross-workspace access If your fabric uses Databricks as the SQL warehouse, you can't select Databricks in an external Source or Target gem. Instead, you must use Table gems, which are limited to the Databricks warehouse defined in the SQL warehouse connection. To work with tables from a different Databricks workspace, use [Delta Sharing](https://docs.databricks.com/aws/en/delta-sharing/). Delta Sharing lets you access data across workspaces without creating additional Databricks connections. Prophecy implements this guardrail to avoid using external connections when the data can be made available in your warehouse. External connections introduce an extra data transfer step, which slows down pipeline execution and adds unnecessary complexity. For best performance, Prophecy always prefers reading and writing directly within the warehouse. ## Reusing and sharing tables After you create a table in Prophecy, you can reuse its configuration across your entire project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To make tables available to other teams, you can share your project as a package in the [Package Hub](/data-analysis/development/extensibility/package-hub/package-hub). Other users will be able to use the shared table configuration, provided they have the necessary permissions in Databricks to access the underlying data. # Prophecy In Memory Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/prophecy-warehouse Read and write tables to Prophecy In Memory This gem runs in . ## Overview In Prophecy, datasets stored in the [SQL Warehouse Connection](/data-analysis/environment/fabrics/prophecy-fabrics) defined in your fabric are accessed using Table gems. Unlike Source and Target gems, Table gems run directly within the data warehouse, eliminating extra orchestration steps and improving performance. Available configurations for Table gems vary based on your SQL warehouse provider. This page explains how to use the Table gem for a [Prophecy In Memory](/data-analysis/environment/fabrics/prophecy-fabrics) fabric. ## Table types The following table types are supported. | Name | Description | Type | | ----- | ------------------------------------------------------------------------------------------------------------- | ---------------- | | Table | Persistent storage of structured data in your SQL warehouse. Faster for frequent queries (indexed). | Source or Target | | View | A virtual table that derives data dynamically from a query. Slower for complex queries (computed at runtime). | Source or Target | | Seed | Small CSV-format files that you can write directly in Prophecy. | **Source only** | ## Create a new table Once you create a Table gem, you can reuse the table throughout your project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To create a table in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Table** from the dropdown. 4. Click on the gem to open the configuration. Click **+ New Table**. The gem configuration will depend on the table type: * [Tables](#tables) * [Views](#views) * [Seeds](#seeds) ## Configure tables ### Source parameters | Parameter | Description | | ---------- | ------------------------------------------------------------------------------------------------------------- | | Location | Specify the table's location using the schema and table name. You cannot add or change the database location. | | Properties | Define or infer schema. Add a description if needed. | | Preview | Load a sample of the data before saving. | ### Target parameters | Parameter | Description | | ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | | Location | Choose the location where the table will be stored. You can create a new table by writing a new table name. You cannot add or change the database location. | | Properties | Define certain properties of the table. The schema cannot be changed for target tables. | | Write Options | Select how you want the data to be written each time you run the pipeline. Wipe and Replace Table only. | The Location tab defines where a table lives and how it is identified within your project. | Field | Description | | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Table alias | A stable, logical identifier for the table that stays constant even if the underlying database, schema, or table name changes. Required when you create the table, and cannot be changed afterward. | | Database | The database containing the table. | | Schema | The schema containing the table. | | Table | The table name. | A **parameter set** selector appears in the upper-right corner of the Location tab: * If the table gem is used inside a pipeline, the selector shows that pipeline's active parameter set automatically. * If you're working on the table outside of a pipeline (for example, from the Project browser), the selector shows **Select Pipeline and Parameter set** until you choose one. You only need to do this if one or more Location fields are set to Advanced mode (see below) and you need their values resolved outside pipeline context. At the bottom of the Location tab, Prophecy shows a live preview of the resolved `database.schema.table` location: the hardcoded values if all fields are in Simple mode, or the resolved parameter values if any field is in Advanced mode. If a value can't be resolved yet β€” for example, because no parameter set is selected β€” Prophecy displays the raw value instead. #### Make a location field dynamic In the default **Simple** mode, each Location field (database, schema, table) takes a fixed value that you type directly. Switch a field to **Advanced** mode to bind it to a project or pipeline parameter instead of a fixed value. Only parameters of type `sql_expression` can be used in Advanced mode. Using a parameter of a different type will cause the table location to fail to resolve. Switching a field from Advanced back to Simple mode clears its current value. Once a field is in Advanced mode, its value depends on which parameter set is active (see the parameter set selector above). This makes it possible to define a table once and reuse it across multiple pipelines, each supplying different values for the parameterized fields via their own parameter sets. To reuse a table you've already created, select it from **Table > \[alias]** in the Project browser. Prophecy does not validate that tables resolved from different parameter sets share the same schema. If your parameter sets point to tables with different schemas, downstream steps in your pipeline may fail or behave unexpectedly. ## Configure views Views are virtual tables recomputed at runtime from a query. ### Source parameters | Parameter | Description | | ---------- | ---------------------------------------------------- | | Location | Enter the database, schema, and table (view) name. | | Properties | Define or infer schema. Add a description if needed. | | Preview | Load data based on the view's underlying query. | ### Target parameters | Parameter | Description | | ---------- | --------------------------------------------------------------------------------------- | | Location | Define the name of the view to be created or replaced. | | Properties | Define certain properties of the table. The schema cannot be changed for target tables. | | Preview | Load a preview of the resulting view. | Every time the pipeline runs, the target is overwritten. This is because the view is recomputed from scratch based on the underlying logic, and any previously materialized results are discarded. No additional write modes are supported. ## Configure seeds Seeds are lightweight CSV datasets defined in your project. Seeds are source-only. | Parameter | Description | | ---------- | ---------------------------------------------------------------------------------------------------------------------- | | Properties | Copy-paste your CSV data and define certain [properties](https://docs.getdbt.com/reference/seed-configs) of the table. | | Preview | Load a preview of your seed in table format. | The CSV data you define is stored in your Prophecy project files and materialized as a table in your data warehouse. This table is created in the [default target schema](/data-analysis/environment/connections/databricks#connection-parameters) specified in your connection. ## Reusing and sharing tables After you create a table in Prophecy, you can reuse its configuration across your entire project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To make tables available to other teams, you can share your project as a package in the [Package Hub](/data-analysis/development/extensibility/package-hub/package-hub). Other users will be able to use the shared table configuration if they have access to your fabric via team membership. # Read from Snowflake tables Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/snowflake-read Configure a Snowflake table, view, or seed as a read source This gem runs in . ## Overview In Prophecy, datasets stored in the [SQL Warehouse Connection](/data-analysis/environment/fabrics/prophecy-fabrics) defined in your fabric are accessed using Table gems. Unlike Source and Target gems, Table gems run directly within the data warehouse, eliminating extra orchestration steps and improving performance. This page explains how to read from a Snowflake table, view, or seed using the Table gem. To write to a Snowflake table or view instead, see [Write to Snowflake tables](/data-analysis/gems/source-target/table/snowflake-write). ## Table types The following table types are supported for Snowflake connections. | Name | Description | Type | | ----- | --------------------------------------------------------------------------------------------------------------- | ---------------- | | Table | Persistent storage of structured data in your SQL warehouse. Optimized for frequent queries and large datasets. | Source or Target | | View | A virtual table that derives data dynamically from a query. Recomputed at runtime. | Source or Target | | Seed | Small CSV-format files that you can write directly in Prophecy. | **Source only** | ## Configure table Once you create a Table gem, you can reuse the table throughout your project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To read from a table in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Table** from the dropdown. 4. Click on the gem to open the configuration. Choose the table, view, or seed you want to read from the list. Seeds configure differently from Tables and Views β€” see [Configure seeds](#configure-seeds) below. Specify the table's location using database, schema, and name. For a **View**, enter the database, schema, and view name here instead. Define or infer schema. Add a description if needed. Load a sample of the data before saving. For a **View**, this loads data based on the view's underlying query rather than stored data. See [Table tests vs. project tests](/data-analysis/development/tests/test-comparison) to decide which approach fits your validation needs. ## Configure seeds Seeds are lightweight CSV datasets defined in your project. Seeds are source-only. | Parameter | Description | | ---------- | ---------------------------------------------------------------------------------------------------------------------- | | Properties | Copy-paste your CSV data and define certain [properties](https://docs.getdbt.com/reference/seed-configs) of the table. | | Preview | Load a preview of your seed in table format. | Seeds are implemented as [dbt seeds](https://docs.getdbt.com/docs/build/seeds) under the hood. The CSV data you define is stored in your Prophecy project files and materialized as a table in your Snowflake data warehouse. This table is created in the default database and schema specified in your Snowflake connection. ## Reusing and sharing tables After you create a table in Prophecy, you can reuse its configuration across your entire project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To make tables available to other teams, you can share your project as a package in the [Package Hub](/data-analysis/development/extensibility/package-hub/package-hub). Other users will be able to use the shared table configuration, provided they have the necessary permissions in Snowflake to access the underlying data. ## Limitations Currently, Snowflake tables do not support: * Case-sensitive identifiers. * Creating new partitioned tables. * Modifying partitioning of existing tables. # Write to Snowflake tables Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/snowflake-write Configure a Snowflake table or view as a write target This gem runs in . ## Overview In Prophecy, datasets stored in the [SQL Warehouse Connection](/data-analysis/environment/fabrics/prophecy-fabrics) defined in your fabric are accessed using Table gems. Unlike Source and Target gems, Table gems run directly within the data warehouse, eliminating extra orchestration steps and improving performance. This page explains how to write to a Snowflake table or view using the Table gem. To read from an existing Snowflake table or view instead, see [Read from Snowflake tables](/data-analysis/gems/source-target/table/snowflake-write) ## Configure table Once you create a Table gem, you can reuse it throughout your project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To write to a table in your pipeline: 1. Open your pipeline in the [Studio](/data-analysis/development/studio/studio). 2. Click on **Source/Target** in the canvas. 3. Select **Table** from the dropdown. 4. Click on the gem to open the configuration. To write to an existing table, select it from the list. To write to a new table, click **+ New Table**. Choose **Table** or **View**. The Location tab defines where a table lives and how it is identified within your project. | Field | Description | | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Table alias | A stable, logical identifier for the table that stays constant even if the underlying database, schema, or table name changes. Required when you create the table, and cannot be changed afterward. | | Database | The database containing the table. | | Schema | The schema containing the table. | | Table | The table name. | A **parameter set** selector appears in the upper-right corner of the Location tab: * If the table gem is used inside a pipeline, the selector shows that pipeline's active parameter set automatically. * If you're working on the table outside of a pipeline (for example, from the Project browser), the selector shows **Select Pipeline and Parameter set** until you choose one. You only need to do this if one or more Location fields are set to Advanced mode (see below) and you need their values resolved outside pipeline context. At the bottom of the Location tab, Prophecy shows a live preview of the resolved `database.schema.table` location: the hardcoded values if all fields are in Simple mode, or the resolved parameter values if any field is in Advanced mode. If a value can't be resolved yet β€” for example, because no parameter set is selected β€” Prophecy displays the raw value instead. #### Make a location field dynamic In the default **Simple** mode, each Location field (database, schema, table) takes a fixed value that you type directly. Switch a field to **Advanced** mode to bind it to a project or pipeline parameter instead of a fixed value. Only parameters of type `sql_expression` can be used in Advanced mode. Using a parameter of a different type will cause the table location to fail to resolve. Switching a field from Advanced back to Simple mode clears its current value. Once a field is in Advanced mode, its value depends on which parameter set is active (see the parameter set selector above). This makes it possible to define a table once and reuse it across multiple pipelines, each supplying different values for the parameterized fields via their own parameter sets. To reuse a table you've already created, select it from **Table > \[alias]** in the Project browser. Prophecy does not validate that tables resolved from different parameter sets share the same schema. If your parameter sets point to tables with different schemas, downstream steps in your pipeline may fail or behave unexpectedly. Map each incoming column to a column on the target table. Prophecy suggests mappings automatically β€” review and adjust them as needed. Any target column left unmapped defaults to null. This is schema **mapping**, not schema definition or inference. Prophecy already knows the target's schema (from an existing table, or as you define one for a new table in the Location step); this step is about reconciling your pipeline's output columns against it. If your BigQuery tables are partitioned or clustered, Prophecy automatically displays the partitioning and clustering information in the Properties tab. You'll be able to view information such as the partition granularity, partitioning data type, clustering columns, and more. You won't be able to edit these properties. You can also set a description for the table and configure generic options, such as skipping execution when the input has zero rows. ### Map schema for existing target tables When you select an existing table as a target, the incoming schema might not match the schema of the target table. Prophecy lets you reconcile these differences in the **Map Schema** section of the Properties tab. For each target column, select the corresponding source column. Prophecy can suggest mappings for unmapped columns, which you can review and select. Any required casts or other transformations are applied so that the incoming data conforms to the target schema. If you want the target table to use the incoming schema instead, click **Overwrite Target Schema**. This replaces the existing target schema with the source schema rather than mapping the incoming columns to it. Schema mapping is available when writing to existing target tables in Snowflake, Databricks, and BigQuery. Select how you want the data to be written each time you run the pipeline. Learn more in [Write strategies](/data-analysis/gems/source-target/table/write/write-options). This step applies to **Table** targets only. If you selected **View** in the type & format step, there's no Write Options tab β€” the gem shows **Preview** instead, even when it's positioned at the end of a pipeline. Views are always fully recomputed and overwritten on each run, so there's no write mode to choose. Add data tests to validate the table after it's written. See [Table tests vs. project tests](/data-analysis/development/tests/test-comparison) to decide which approach fits your validation needs. ## Reusing and sharing tables After you create a table in Prophecy, you can reuse its configuration across your entire project. All created tables appear in the [Project](/data-analysis/development/studio/studio) tab in the left sidebar. To make tables available to other teams, you can share your project as a package in the [Package Hub](/data-analysis/development/extensibility/package-hub/package-hub). Other users will be able to use the shared table configuration, provided they have the necessary permissions in Snowflake to access the underlying data. ## Limitations Currently, Snowflake tables do not support: * Case-sensitive identifiers. * Creating new partitioned tables. * Modifying partitioning of existing tables. # Upload files gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/upload-files Upload files to your data warehouse from the visual canvas You can add a source table to your primary SQL warehouse by uploading a file directly onto the visual canvas. This gives you greater control over your data and how you incorporate it into your model transformation. When you upload your file through Prophecy, it's added directly to your SQL warehouse as a [table](/data-analysis/gems/source-target/source-target). The recommended maximum file size is 100Β MB. You can upload the following file formats: * CSV * Excel (XLS, XLSX) * JSON * Parquet * XML ## File upload To upload your file: 1. Open a pipeline. 2. Drag and drop your file onto the canvas. This opens a new Source gem. Upload file by dragging and dropping ## Type and format Once the Source gem opens, you can review the type and format of the file. 1. Optional: Replace or delete your uploaded file. 2. Optional: Change the type and format of the file. 3. Click **Next**. Select your file type and format ## Location Now, set the location where the table will be stored in the primary SQL warehouse configured in your fabric. 1. Choose the database. 2. Choose the schema. 3. Choose the table. You can either select an existing table or create a new one. Select the table location If you select an existing table, Prophecy deletes and recreates the table with your uploaded file. ## Properties You can configure the table properties before completing the file upload. 1. Review the file's options. Depending on the file type and format, common defaults are already chosen for you. 2. Optional: Modify the options. For example, you can change the header row by selecting **First row is header**. 3. Optional: If you made any changes to the options, click **Infer Schema**. This will update the schema according to the defined options. Configure the table properties Excel files (XLS, XLSX) expose additional options in the Properties panel: | Field | Purpose | | -------------------------- | ------------------------------------------------------------------------------------------------ | | **Read mode** | How the workbook is read: a single sheet, or a union of multiple sheets combined into one table. | | **Sheet Name** | Which sheet to read from. When reading a union of multiple sheets, select each sheet to include. | | **Enter Cell Range** | Restrict extraction to a specific range, for example `A1:`. | | **Header** | Treat the first row of the range as column headers. | | **Allow Undefined Rows** | Include rows that fall outside the detected schema. | | **Allow Incomplete Rows** | Include rows missing values in some columns. | | **Enable Schema Merging** | Merge schemas across sheets or tables where they differ. | | **Ignore Cell Formatting** | Read cell values without applying Excel display formatting. | After you configure these options, click **Infer Schema** to populate the column list. Column names, types (such as `String` or `Bigint`), and other metadata are editable. ### Extract multiple tables from one workbook You can mark multiple regions of an Excel workbook as separate tables. Each marked table becomes its own output port on the gem, and you can: * Edit the schema for each table independently. * Set drift-handling behavior per table. * Unmark or delete a table you no longer need. ## Preview The preview step shows your table data and gives you the option to download it. 1. To generate the preview, click **Load**. This loads the data. 2. Check that your preview looks correct and click **Done**. 3. If you selected a table to write your uploaded file to, you'll need to confirm the upload in the pop-up window by clicking **Proceed**. Preview the table For Excel files, the preview grid loads additional rows and columns as you scroll, rather than fetching the entire sheet at once. Merged cells and blank rows in the source sheet are handled automatically. The new table is now in your environment and available in Source/Target gems. You can upload another file or start working with your new source gem. ## Connect extracted tables to your pipeline Once a table is extracted, connect it into the canvas like any other dataset using the standard add-connection menu. Hover over an output port to see a schema hint before connecting it. # Append Rows Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/write/append Always add incoming rows to the existing table When you choose the **Append Rows** option, Prophecy adds new rows to the existing table without modifying existing data. No deduplication is performed, so you may end up with duplicate records. This strategy is best used when unique keys aren't required. For key-based updates, use one of the Merge options instead. ## Example: Append daily sales A sales table receives a new batch of orders each day. The incoming dataset is appended to the existing table so that all historical orders are preserved.
**Existing target table** | ORDER\_ID | PRODUCT | SALE\_DATE | | --------- | ------- | ---------- | | 101 | Laptop | 2024-01-14 | | 102 | Mouse | 2024-01-14 | **Incoming dataset** | ORDER\_ID | PRODUCT | SALE\_DATE | | --------- | -------- | ---------- | | 301 | Monitor | 2024-01-15 | | 302 | Keyboard | 2024-01-15 | **Resulting target table** | ORDER\_ID | PRODUCT | SALE\_DATE | | --------- | -------- | ---------- | | 101 | Laptop | 2024-01-14 | | 102 | Mouse | 2024-01-14 | | 301 | Monitor | 2024-01-15 | | 302 | Keyboard | 2024-01-15 | Prophecy appends every row from the incoming dataset to the target table without modifying or removing existing rows. In this example, orders `301` and `302` are added while orders `101` and `102` remain unchanged.
# Append Unique Rows Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/write/append-unique Insert only new records, leaving matching records unchanged When you choose the **Append Unique Rows** option, if a row's unique key already exists in the target table, the incoming row is ignored and the existing row is left unchanged. Only rows with a unique key that doesn't yet exist in the target are inserted. Unlike [Upsert Row](/data-analysis/gems/source-target/table/write/upsert), this mode never modifies existing records; it only inserts new ones. ## Parameters | Parameter | Description | | -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | | Unique Key | Column(s) used to find records that already exist in the target table. Matching records are left unchanged; only the remaining records are inserted. | | Use Predicate | Limits which target rows are checked for duplicates. Keys outside the predicate range may be re-inserted. | | Use a condition to filter data or incremental runs | Enables applying conditions for filtering the incoming data into the table. | Merge Columns, Exclude Columns, and On Schema Change don't apply to this write mode. If you switch to Append Unique Rows from another merge approach, any values set for those options are cleared. ## Example: Append new orders only An order table uses `ORDER_ID` as the unique key. The incoming dataset includes a status update for an existing order and one new order.
**Existing target table** | ORDER\_ID | STATUS | AMOUNT | | --------- | ------- | ------ | | 101 | Shipped | 250 | | 102 | Pending | 100 | ### Incoming dataset | ORDER\_ID | STATUS | AMOUNT | | --------- | --------- | ------ | | 101 | Cancelled | 250 | | 103 | Pending | 75 | ### Resulting target table | ORDER\_ID | STATUS | AMOUNT | | --------- | ------- | ------ | | 101 | Shipped | 250 | | 102 | Pending | 100 | | 103 | Pending | 75 | Because `ORDER_ID` `101` already exists in the target table, its incoming row is ignored, so the existing `Shipped` status is left unchanged. `ORDER_ID` `103` doesn't exist in the target, so it's inserted as a new row. `ORDER_ID` `102` has no matching row in the incoming dataset, so it's left as-is.
# Delete and Insert Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/write/delete-insert Delete matching rows and insert new rows When you choose the **Delete and Insert** option, Prophecy proceeds as follows: * For rows in the target table that have keys that match rows in the incoming dataset, deletes matching rows from the existing table and replaces them. * For incoming rows with new keys, inserts these normally. This ensures that updated records are fully replaced instead of partially updated. This strategy helps when your unique key is not truly unique (as when multiple rows per key need to be fully refreshed). ## Parameters | Parameter | Description | | -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Unique Key | Column(s) used to match existing records in the target dataset. | | Use Predicate | Lets you add conditions that specify when to apply the merge. | | Use a condition to filter data or incremental runs | Enables applying conditions for filtering the incoming data into the table. | | (Advanced) On Schema Change | Specifies how schema changes should be handled during the merge process.
  • ignore: Newly added columns will not be written to the model. This is the default option.
  • fail: Triggers an error message when the source and target schemas diverge.
  • append\_new\_columns: Append new columns to the existing table.
  • sync\_all\_columns: Adds any new columns to the existing table, and removes any columns that are now missing. Includes data type changes. This option uses the output of the previous gem.
| ## Example: Refresh order line items An order table uses `order_id` as the unique key. When a corrected order arrives, Prophecy deletes all existing rows with the matching `order_id` and inserts the incoming rows. Orders with new IDs are inserted normally.
**Existing target table** | ORDER\_ID | LINE\_ITEM | PRODUCT | | --------- | ---------- | -------- | | 101 | 1 | Laptop | | 101 | 2 | Mouse | | 102 | 1 | Keyboard | **Incoming dataset** | ORDER\_ID | LINE\_ITEM | PRODUCT | | --------- | ---------- | ------- | | 101 | 1 | Laptop | | 101 | 2 | Monitor | | 101 | 3 | Mouse | | 103 | 1 | Dock | **Resulting target table** | ORDER\_ID | LINE\_ITEM | PRODUCT | | --------- | ---------- | -------- | | 101 | 1 | Laptop | | 101 | 2 | Monitor | | 101 | 3 | Mouse | | 102 | 1 | Keyboard | | 103 | 1 | Dock | Because `order_id` is the unique key, Prophecy deletes all existing rows with matching keys before inserting the incoming rows. In this example, both existing rows for order `101` are removed and replaced with the three incoming rows. Order `102` remains unchanged because it has no matching incoming rows, and order `103` is inserted as a new order.
# SCD2 Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/write/scd2 Track historical changes by adding new rows The **Merge - SCD2** write mode tracks historical changes by adding new rows instead of updating existing ones. This lets you preserve a complete history, since every change generates a new entry rather than erasing old information. To be specific: * New rows are added for incoming records with unique keys that don't exist in the target table. * For records matching existing unique keys, a new row is added only when the incoming data differs from the existing record. * New rows are assigned a start date and a `null` end date (indicating it's currently valid). * If the unique key of the new record matched an existing record, the existing row is assigned an end date to mark when it stopped being valid. * This creates a complete timeline showing how data evolved over time. ## Parameters | Parameter | Description | | ----------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Unique Key | Column(s) used to match existing records in the target dataset. In many cases, the unique key will be equivalent to your table's primary key, if applicable. | | Invalidate deleted rows | When enabled, records that match deleted rows will be marked as no longer valid. | | Determine new records by checking timestamp column | Recognizes new records by the time from the **Updated at** column that you define. | | Determine new records by looking for differences in column values | Recognizes new records based on a change of values in one or more specified columns. | ## Example: Route assignment history A vehicle assignment table tracks which route each vehicle is assigned to over time. When a vehicle is assigned to a new route, Prophecy preserves the previous assignment and creates a new current record. In this example, Determine new records by checking timestamp column is enabled, and `assigned_at` is the timestamp column.
**Existing target table** | VEHICLE\_ID | ROUTE | ASSIGNED\_AT | valid\_from | valid\_to | | ----------- | ------- | ------------ | ----------- | --------- | | 101 | Route A | 2024-01-01 | 2024-01-01 | NULL | **Incoming data set** | VEHICLE\_ID | ROUTE | ASSIGNED\_AT | | ----------- | ------- | ------------ | | 101 | Route B | 2024-01-15 | **Updated target table** | VEHICLE\_ID | ROUTE | ASSIGNED\_AT | valid\_from | valid\_to | | ----------- | ------- | ------------ | ----------- | ---------- | | 101 | Route A | 2024-01-01 | 2024-01-01 | 2024-01-15 | | 101 | Route B | 2024-01-15 | 2024-01-15 | NULL | Because the incoming row matches an existing `vehicle_id` but has a newer `assigned_at` value, Prophecy preserves the existing row by assigning it a `valid_to` date and inserts a new current row with a `null` `valid_to` value. This creates a complete history of the vehicle's route assignments rather than overwriting the previous assignment.
## Parameters (PySpark only) Private Preview When PySpark is the project language, the SCD2 write mode includes the following configuration options. | Parameter | Description | Default | | --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | | Key Columns | Column(s) used to identify unique records. These columns remain constant across all versions of a record and are used to match incoming data with existing rows. | None | | Historic Columns | Column(s) that change over time and require historical tracking. When values in these columns differ between incoming and existing records, a new row is created. | None | | From Time Column | Column that stores the start time indicating when a row becomes valid. This timestamp marks the beginning of the validity period for that record version. | None | | To Time Column | Column that stores the end time indicating when a row stops being valid. Set to `null` for the current active version of a record. | None | | Create Min Max Flags | When enabled, creates flag columns to identify the first (minimum) and last (maximum) entries for each unique key. Useful for quickly identifying historical boundaries. | false | | Name of the column used as min/old-value flag | Column name for the minimum flag. Set to `true` (or `1` depending on flag values) for the first historical entry of each key. | min\_flag | | Name of the column used as max/latest flag | Column name for the maximum flag. Set to `true` (or `1` depending on flag values) for the most recent active entry of each key. | max\_flag | | Flag values | Format for min and max flag values. Options are `true/false` (boolean) or `0/1` (numeric). | true/false | | Enable Soft Delete | When enabled, deleted records from the source are treated as Type 2 changes. The target table updates the deleted record's end time instead of removing the row, preserving deletion history. | false | # Upsert Row Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/write/upsert Update existing rows or insert new rows When you choose the **Upsert Row** option, if a row with the same key exists, it is updated. Otherwise, a new row is inserted. You can also limit updates to specific columns, so only selected values are changed in matching rows. ## Parameters | Parameter | Description | | -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Unique Key | Column(s) used to match existing records in the target dataset. In many cases, the unique key will be equivalent to your table's primary key, if applicable. | | Use Predicate | Lets you add conditions that specify when to apply the merge. | | Use a condition to filter data or incremental runs | Enables applying conditions for filtering the incoming data into the table. | | Merge Columns | Specifies which columns to update during the merge. If empty, the merge includes all columns. | | Exclude Columns | Defines columns that should be excluded from the merge operation. | | (Advanced) On Schema Change | Specifies how schema changes should be handled during the merge process.
  • ignore: Newly added columns will not be written to the model. This is the default option.
  • fail: Triggers an error message when the source and target schemas diverge.
  • append\_new\_columns: Append new columns to the existing table.
  • sync\_all\_columns: Adds any new columns to the existing table, and removes any columns that are now missing. Includes data type changes. This option uses the output of the previous gem.
| ## Example: Upsert vehicle types A vehicle registry uses `vehicle_id` as the merge key. The merge column is set to `type`, so matching vehicles have their type updated. Vehicles that do not exist in the target table are inserted as new rows.
**Existing target table** | VEHICLE\_ID | TYPE | REGISTERED\_AT | | ----------- | ----- | -------------- | | 101 | Bus | 2023-12-01 | | 102 | Train | 2023-12-02 | | 103 | Bus | 2023-12-03 | **Incoming dataset** | VEHICLE\_ID | TYPE | | ----------- | ---- | | 101 | Tram | | 102 | Bus | | 104 | Bus | **Resulting target table** | VEHICLE\_ID | TYPE | REGISTERED\_AT | | ----------- | ---- | -------------- | | 101 | Tram | 2023-12-01 | | 102 | Bus | 2023-12-02 | | 103 | Bus | 2023-12-03 | | 104 | Bus | null | Because the merge column is set to `TYPE`, Prophecy updates only that column for matching rows. Columns that are not included in the merge remain unchanged. Rows that do not match the merge key are inserted as new rows. In this example, vehicles `101` and `102` have updated `TYPE` values, vehicle `103` remains unchanged, and vehicle `104` is inserted with a `null` value for `REGISTERED_AT` because that column was not included in the incoming dataset.
# Wipe and Replace Partitions Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/write/wipe-replace-partitions Replace all the rows in a partition When you choose the **Wipe and Replace Partitions** option, Prophecy replaces entire partitions in the target table. Only partitions containing updated data will be overwritten; other partitions will not be modified. ## Parameters | Parameter | Description | | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Partition by | Defines the partitions of the target table.
  • Databricks: Each unique value of the partition column corresponds to a partition. You cannot change the granularity of the partitions.
  • BigQuery: You must [manually define the granularity](#define-partition-granularity-bigquery-only) of your partitions. BigQuery does not automatically infer how to write the partitions.
| | (Advanced) On Schema Change | Specifies how schema changes should be handled during the merge process.
  • ignore: Newly added columns will not be written to the model. This is the default option.
  • fail: Triggers an error message when the source and target schemas diverge.
  • append\_new\_columns: Append new columns to the existing table.
  • sync\_all\_columns: Adds any new columns to the existing table, and removes any columns that are now missing. Includes data type changes. This option uses the output of the previous gem.
| ## Example: Replace one day's sales A sales table is partitioned by the `SALE_DATE` column. A corrected dataset arrives for `2024-01-15`, so only that partition is replaced. All other partitions remain unchanged.
**Existing target table** | ORDER\_ID | PRODUCT | SALE\_DATE | | --------- | ------- | ---------- | | 101 | Laptop | 2024-01-14 | | 201 | Mouse | 2024-01-15 | | 202 | Monitor | 2024-01-15 | **Incoming dataset** | ORDER\_ID | PRODUCT | SALE\_DATE | | --------- | -------- | ---------- | | 301 | Mouse | 2024-01-15 | | 302 | Keyboard | 2024-01-15 | **Resulting target table** | ORDER\_ID | PRODUCT | SALE\_DATE | | --------- | -------- | ---------- | | 101 | Laptop | 2024-01-14 | | 301 | Mouse | 2024-01-15 | | 302 | Keyboard | 2024-01-15 | Because this write mode replaces entire partitions, Prophecy removes all existing rows from any partition represented in the incoming dataset before writing the new rows. Partitions that are not represented in the incoming dataset are left unchanged. In this example, only the `2024-01-15` partition is replaced, while the `2024-01-14` partition remains unchanged.
## Define partition granularity (BigQuery only) The following partitioning parameters allow you to define the partition granularity for this operation. | Parameter | Description | | ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Column Name | The name of the column used for partitioning the target table. | | Data Type | The data type of the partition column.
Supported types: `timestamp`, `date`, `datetime`, and `int64`. | | Partition By granularity | Applicable only to `timestamp`, `date`, or `datetime` data type.
Defines the time-based partition granularity: `hour`, `day`, `month`, or `year`. | | Partition Range | Applicable only to `int64` data type.
Specify a numeric range for partitioning using a **start**, **end**, and **interval** value (e.g., start=`0`, end=`1000`, interval=`10`).
You must define an interval value so that Prophecy knows at what intervals to create the partitions. | # Wipe and Replace Rows Per Predicate Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/write/wipe-replace-predicate Replace all the rows in a partition that match a predicate When you choose the **Wipe and Replace Rows Per Predicate** option, Prophecy updates records that match conditions defined in the **Use Predicate** parameter. Instead of updating individual rows, it replaces the entire partition of the table that matches the given predicate. The predicate defines which rows in the target table should be fully replaced with the incoming data. ## Parameters | Parameter | Description | | -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Use Predicate | Lets you add conditions that specify when to apply the merge. | | Use a condition to filter data or incremental runs | Enables applying conditions for filtering the incoming data into the table. | | (Advanced) On Schema Change | Specifies how schema changes should be handled during the merge process.
  • ignore: Newly added columns will not be written to the model. This is the default option.
  • fail: Triggers an error message when the source and target schemas diverge.
  • append\_new\_columns: Append new columns to the existing table.
  • sync\_all\_columns: Adds any new columns to the existing table, and removes any columns that are now missing. Includes data type changes. This option uses the output of the previous gem.
| ## Example: Replace sales from January 5 onward A sales table uses the predicate `sale_date >= '2024-01-05'`. When a corrected dataset arrives, Prophecy removes all rows that satisfy the predicate and replaces them with the incoming dataset. Rows that do not satisfy the predicate remain unchanged. **Predicate**: `sale_date >= '2024-01-05'`
**Existing target table** | ORDER\_ID | PRODUCT | SALE\_DATE | STATUS | | --------- | -------- | ---------- | ------ | | 101 | Laptop | 2023-12-20 | Open | | 201 | Mouse | 2024-01-05 | Open | | 202 | Monitor | 2024-01-10 | Open | | 203 | Keyboard | 2024-01-15 | Open | **Incoming dataset** | ORDER\_ID | PRODUCT | SALE\_DATE | STATUS | | --------- | ------- | ---------- | ------ | | 201 | Mouse | 2024-01-05 | Closed | | 202 | Monitor | 2024-01-10 | Closed | **Resulting target table** | ORDER\_ID | PRODUCT | SALE\_DATE | STATUS | | --------- | ------- | ---------- | ------ | | 101 | Laptop | 2023-12-20 | Open | | 201 | Mouse | 2024-01-05 | Closed | | 202 | Monitor | 2024-01-10 | Closed | Because this write mode replaces all rows that satisfy the predicate, Prophecy removes every existing row that matches the predicate before writing the incoming dataset. Rows that do not satisfy the predicate remain unchanged. In this example, the row for `2024-01-15` (`ORDER_ID` `203`) is removed because it matches the predicate but is not included in the incoming dataset, while the row from `2023-12-20` remains unchanged.
# Wipe and Replace Table Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/write/wipe-replace-table Replace all the rows in the target table On each run, Prophecy removes all existing rows from the target table and writes the incoming dataset. The target table always reflects the latest state of the source data. ## Example: Replace product list A product catalog is refreshed each day by replacing the existing table with the latest product list.
**Existing target table** | PRODUCT\_ID | NAME | QUANTITY | | ----------- | -------- | -------- | | 101 | Laptop | 25 | | 102 | Mouse | 84 | | 103 | Keyboard | 40 | **Incoming data set** | PRODUCT\_ID | NAME | QUANTITY | | ----------- | ------- | -------- | | 101 | Laptop | 21 | | 104 | Monitor | 17 | **Resulting target table** | PRODUCT\_ID | NAME | QUANTITY | | ----------- | ------- | -------- | | 101 | Laptop | 21 | | 104 | Monitor | 17 | Because this write mode replaces the entire table, any rows that are not present in the incoming dataset are removed. Prophecy does not compare existing and incoming rows or merge individual records. In this example, products `102` and `103` are removed because they are not included in the latest product list.
## Partition the target table (BigQuery only) The following partitioning parameters are available for the **Wipe and Replace Table** write mode on BigQuery. | Parameter | Description | | ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Column Name | The name of the column used for partitioning the target table. | | Data Type | The data type of the partition column.
Supported types: `timestamp`, `date`, `datetime`, and `int64`. | | Partition By granularity | Applicable only to `timestamp`, `date`, or `datetime` data type.
Defines the time-based partition granularity: `hour`, `day`, `month`, or `year`. | | Partition Range | Applicable only to `int64` data type.
Specify a numeric range for partitioning using a **start**, **end**, and **interval** value (e.g., start=`0`, end=`1000`, interval=`10`).
You must define an interval value so that Prophecy knows at what intervals to create the partitions. | Only BigQuery tables can be partitioned at the table level. To learn more about partitioning, jump to [Partitioning](#partitioning). # Write options Source: https://docs.prophecy.ai/data-analysis/gems/source-target/table/write/write-options Choose from multiple write strategies including append, merge, and replace Prophecy provides a set of write strategies that determine how you will store your processed data and handle changes to the data over time. This page describes each strategy so you can choose the best one for your use case. You will configure the write strategy in the **Write Options** tab of a target table. ## Write modes matrix The following table describes the write modes that Prophecy supports by SQL warehouse and gem type. | Write mode | Databricks table | Databricks model | BigQuery table | BigQuery model | Snowflake model | | ------------------------------------------------------------------------------------------------------------------- | ---------------- | ---------------- | -------------- | -------------- | --------------- | | [Wipe and Replace Table (Default)](/data-analysis/gems/source-target/table/write/wipe-replace-table) | βœ” | βœ” | βœ” | βœ” | βœ” | | [Append Row](/data-analysis/gems/source-target/table/write/append) | βœ” | βœ” | βœ” | | βœ” | | [Merge - Upsert Row](/data-analysis/gems/source-target/table/write/upsert) | βœ” | βœ” | βœ” | βœ” | βœ” | | [Merge - Append Unique Rows](/data-analysis/gems/source-target/table/write/append-unique) | βœ” | | | | | | [Merge - Wipe and Replace Partitions](/data-analysis/gems/source-target/table/write/wipe-replace-partitions) | βœ” | βœ” | βœ” | βœ” | | | [Merge - SCD2](/data-analysis/gems/source-target/table/write/scd2) | βœ” | βœ” | βœ” | βœ” | βœ” | | [Merge - Wipe and Replace Rows Per Predicate](/data-analysis/gems/source-target/table/write/wipe-replace-predicate) | βœ” | βœ” | | | | | [Merge - Delete and Insert](/data-analysis/gems/source-target/table/write/delete-insert) | | | | | βœ” | ## How write modes work Prophecy simplifies data transformation by providing intuitive write mode options that abstract away the complexity of underlying SQL operations. Behind the scenes, Prophecy generates dbt models that implement these write strategies using SQL warehouse-specific commands. When you select a write mode in a Table gem, Prophecy automatically generates the appropriate dbt configuration and SQL logic. This means you can focus on your data transformation logic rather than learning dbt's materialization strategies or writing complex SQL merge statements. To understand exactly what happens when Prophecy runs these write operations, switch to the **Code** view of your project and inspect the generated dbt model files. These files contain the SQL statements and dbt configuration (like `materialized: 'incremental'`) that dbt uses to execute the write operation. To learn more about the specific configuration options available for each SQL warehouse, visit the dbt documentation links below. * [BigQuery configurations](https://docs.getdbt.com/reference/resource-configs/bigquery-configs) * [Databricks configurations](https://docs.getdbt.com/reference/resource-configs/databricks-configs) * [Snowflake configurations](https://docs.getdbt.com/reference/resource-configs/snowflake-configs) ## Partitioning Depending on the SQL warehouse you use to write tables, partitioning can have different behavior. Let's examine the differences between partitioning in Databricks, Google BigQuery, and Snowflake. In Databricks, partitioning is a write strategy. Databricks organizes data in folders by column values. Partitioning only makes sense when you're using the **Wipe and Replace Partitions** write mode because it allows you to overwrite specific directories (partitions) without rewriting the whole table. For the **Wipe and Replace Table** option, the table is dropped and completely recreated. Partitioning doesn't add any runtime benefit here, so this is not an option for Databricks. In Google BigQuery, partitioning is a table property. Partitioning is defined at the table schema level (time, integer range, or column value). Because it is a part of the table architecture, the physical storage in BigQuery is optimized by partitioning automatically. Once a table is partitioned, every write to that table, full or incremental, respects the partitioning. That means even when you drop and create an entirely new table, Google BigQuery creates the table with partitions in an optimized way. In Snowflake, tables are automatically divided into **micro-partitions** as data is loaded. Unlike BigQuery and Databricks, you don't define partition columns or manage individual partitions when writing a table. Snowflake automatically manages the physical organization of micro-partitions. For large tables, you can optionally define clustering keys to influence how data is organized and improve query pruning, but clustering is separate from the write strategy used to populate the table. As a result, Snowflake write modes don't require you to configure partitioning when writing data. ## Troubleshooting This happens when the incoming and existing schemas don't align. To solve this, use the "On Schema Change" setting to set behavior or ensure schema compatibility. This happens mainly when using the **Append Rows** write mode. To solve this, consider using one of the merge modes instead of append. # Buffer Source: https://docs.prophecy.ai/data-analysis/gems/spatial/buffer Expand or contracts the boundaries of a polygon or line This gem runs in . ## Overview Use the Buffer gem to take any polygon or line and expand or contract its boundaries. This can be useful for spatial analysis tasks like creating safety zones around hazardous areas, expanding service coverage areas, and analyzing proximity impacts. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `ProphecyDatabricksSqlSpatial` version 0.0.3 or higher to your project. * Run Databricks Runtime 17.1 or higher. (This requirement is due to the use of Databricks' [ST geospatial functions](https://docs.databricks.com/aws/en/sql/language-manual/sql-ref-st-geospatial-functions).) Other SQL warehouse providers are not supported for this gem. ## Input and Output The Buffer gem accepts the following inputs and output. | Port | Description | | ------- | ------------------------------------------------------------------------------------------------------------------------- | | **in0** | Input dataset containing the source points in WKT format for which you want to find the nearest points. | | **out** | Output dataset that contains two columns: `input` with the original geometry, and `output` with the transformed geometry. | Input geometries must be in Well-known Text ([WKT](https://en.wikipedia.org/wiki/Well-known_text_representation_of_geometry)) geometric format. Use the [PolyBuild](/data-analysis/gems/spatial/polybuild) gem to create lines and polygons in this format from latitude and longitude coordinates. ## Parameters Configure the Buffer gem using the following parameters. | Parameter | Description | | --------------- | ------------------------------------------------------------------------------------------------------------------------------- | | Geometry column | Column containing the polygon or line you want to expand or contract. | | Distance | Amount of distance to expand or contract each geometry. Use negative distances to create inward buffers (shrinking geometries). | | Units | Unit of measurement for the distance you defined. | ## Example Let's say you're working with a transportation dataset and need to create safety corridors around major highways. You have highway routes as polylines and want to create 3 mile buffer zones on both sides of each road for noise impact analysis. 1. Add a Buffer gem to your pipeline canvas. 2. Attach an input that includes the highway routes as polylines in a `routes` column. 3. Open the gem configuration interface. 4. For **Geometry column**, select the `routes` column from the input table. 5. For **Distance**, input `3`. 6. For **Units**, select **Miles**. 7. Save and run the gem. ### Result The output will contain both the original routes and your transformed highway routes. # CreatePoint Source: https://docs.prophecy.ai/data-analysis/gems/spatial/create-point Create geographic points with longitude and latitude coordinates This gem runs in . ## Overview Use the CreatePoint gem to convert longitude and latitude coordinates into geographic points. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `ProphecyDatabricksSqlSpatial` version 0.0.1 or higher to your project. ## Input and Output The CreatePoint gem accepts the following input and output. | Port | Description | | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | **in0** | Source dataset containing pairs of columns with longitude and latitude coordinates. | | **out** | Output dataset containing one or more new columns with points in Well-known Text ([WKT](https://en.wikipedia.org/wiki/Well-known_text_representation_of_geometry)) geometric format. | ## Parameters The CreatePoint gem accepts longitude and latitude columns as parameters. In the **Create Spatial Points** section, you need to **Click to Add a Point**. For each point you add, fill in the following information. | Parameter | Description | | --------------------- | ---------------------------------------------------------------- | | Longitude Column Name | Input column that contains longitude values. | | Latitude Column Name | Input column that contains latitude values. | | Target Column Name | Column in the gem output that will contain resulting geo points. | ## Example Assume you have the following airline route table. | `start_city` | `start_lat` | `start_long` | `destination_city` | `destination_lat` | `destination_long` | | ------------ | ----------- | ------------ | ------------------ | ----------------- | ------------------ | | New York | 40.7128 | -74.0060 | Los Angeles | 34.0522 | -118.2437 | | London | 51.5074 | -0.1278 | Paris | 48.8566 | 2.3522 | | Tokyo | 35.6895 | 139.6917 | Sydney | -33.8688 | 151.2093 | | Toronto | 43.6511 | -79.3470 | Chicago | 41.8781 | -87.6298 | | Dubai | 25.2760 | 55.2962 | Mumbai | 19.0760 | 72.8777 | Scroll horizontally to view the full table. To convert the start and destination coordinates into geographic points: 1. Create the first column pairing: 1. Click **Add a Point**. 2. For **Longitude Column Name**, select the `start_long` column. 3. For **Latitude Column Name**, select the `start_lat` column. 4. For the **Target Column Name**, type `source_point`. 2. Create another column pairing: 1. Click **Add a Point**. 2. For **Longitude Column Name**, select the `destination_long` column. 3. For **Latitude Column Name**, select the `destination_lat` column. 4. For the **Target Column Name**, type `dest_point`. ### Result The CreatePoint gem will produce the following output with two new columns: `source_point` and `dest_point`. | `start_city` | `start_lat` | `start_long` | `destination_city` | `destination_lat` | `destination_long` | `source_point` | `dest_point` | | ------------ | ----------- | ------------ | ------------------ | ----------------- | ------------------ | ------------------------ | ------------------------- | | New York | 40.7128 | -74.0060 | Los Angeles | 34.0522 | -118.2437 | POINT (-74.0060 40.7128) | POINT (-118.2437 34.0522) | | London | 51.5074 | -0.1278 | Paris | 48.8566 | 2.3522 | POINT (-0.1278 51.5074) | POINT (2.3522 48.8566) | | Tokyo | 35.6895 | 139.6917 | Sydney | -33.8688 | 151.2093 | POINT (139.6917 35.6895) | POINT (151.2093 -33.8688) | | Toronto | 43.6511 | -79.3470 | Chicago | 41.8781 | -87.6298 | POINT (-79.3470 43.6511) | POINT (-87.6298 41.8781) | | Dubai | 25.2760 | 55.2962 | Mumbai | 19.0760 | 72.8777 | POINT (55.2962 25.2760) | POINT (72.8777 19.0760) | Scroll horizontally to view the full table. # Distance Source: https://docs.prophecy.ai/data-analysis/gems/spatial/distance Calculate the distance between two points This gem runs in . ## Overview Use the Distance gem to calculate the distance between two geographic points. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `ProphecyDatabricksSqlSpatial` version 0.0.1 or higher to your project. ## Parameters Configure the Distance gem using the following parameters. Geographic points must be in Well-known Text ([WKT](https://en.wikipedia.org/wiki/Well-known_text_representation_of_geometry)) format. Use the [CreatePoint](/data-analysis/gems/spatial/create-point) gem to convert longitude and latitude coordinates to WKT format. ### Spatial Object Fields Use these parameters to specify the columns containing the source and destination geographic points for distance calculation. | Parameter | Description | | ------------------ | ------------------------------------------------ | | Source Type | Format of the source column. | | Source Column | Column that contains the source geo points. | | Destination Type | Format of the destination column. | | Destination Column | Column that contains the destination geo points. | ### Select Output Options Use the checkboxes defined below to choose which output columns the Distance gem should generate. | Checkbox | Description | | --------------------------- | ----------------------------------------------------------------------------------------------------- | | Output Distance | Return a column that includes the distance between points in a specified unit of distance | | Output Cardinal Direction | Return a column that includes the cardinal direction from the source point to the destination point | | Output Direction in Degrees | Return a column that includes the direction in degrees from the source point to the destination point | You can select zero, one, or multiple checkboxes. All checkboxes are disabled by default. ## Example Assume you have the following airline route table, and you would like to calculate the distance between start and destination cities. | `start` | `destination` | `src_point` | `dst_point` | | -------- | ------------- | ------------------------ | ------------------------- | | New York | Los Angeles | POINT (-74.0060 40.7128) | POINT (-118.2437 34.0522) | | London | Paris | POINT (-0.1278 51.5074) | POINT (2.3522 48.8566) | | Tokyo | Sydney | POINT (139.6917 35.6895) | POINT (151.2093 -33.8688) | | Toronto | Chicago | POINT (-79.3470 43.6511) | POINT (-87.6298 41.8781) | | Dubai | Mumbai | POINT (55.2962 25.2760) | POINT (72.8777 19.0760) | To find the distance and direction between cities, set the following gem configurations. 1. Set source and destination: 1. Set **Source Type** to **Point**. 2. Set **Source Column** to `start`. 3. Set **Destination Type** to **Point**. 4. Set **Destination Column** to `destination`. 2. Set output options: 1. Select the **Output Distance** checkbox. 2. From the **Units** dropdown, select **Kilometers**. 3. Select the **Output Cardinal Direction** checkbox. 4. Select the **Output Direction in Degrees** checkbox. 5. Run the gem. ### Result The resulting table will have three new columns: `distance_kilometers`, `cardinal_direction`, and `direction_degrees`. | `start` | `destination` | `src_point` | `dst_point` | `distance_kilometers` | `cardinal_direction` | `direction_degrees` | | -------- | ------------- | ------------------------ | ------------------------- | --------------------- | -------------------- | ------------------- | | New York | Los Angeles | POINT (-74.0060 40.7128) | POINT (-118.2437 34.0522) | 3944.42 | W | 273.7 | | London | Paris | POINT (-0.1278 51.5074) | POINT (2.3522 48.8566) | 343.92 | SE | 148.1 | | Tokyo | Sydney | POINT (139.6917 35.6895) | POINT (151.2093 -33.8688) | 7792.96 | S | 169.9 | | Toronto | Chicago | POINT (-79.3470 43.6511) | POINT (-87.6298 41.8781) | 705.64 | W | 256.6 | | Dubai | Mumbai | POINT (55.2962 25.2760) | POINT (72.8777 19.0760) | 1936.68 | E | 107.3 | Scroll horizontally to view the full table. # Heatmap Source: https://docs.prophecy.ai/data-analysis/gems/spatial/heatmap Generate spatial heatmaps from geo point data using hexagons This gem runs in . ## Overview The HeatMap gem transforms latitude and longitude point data into a spatial heatmap using hexagonal tiling. It's useful for identifying clusters of activity or density within a geographic region. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `ProphecyDatabricksSqlSpatial` version 0.0.3 or higher to your project. ## Input and Output The HeatMap gem accepts the following inputs and output. | Port | Description | | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **in0** | Input dataset containing pairs of columns with longitude and latitude coordinates. | | **out** | Output dataset with two columns:
  • The `density` column describes the number of points (or total heat) in the hexagon.
  • The `geometry_wkt` column includes the hexagon boundary in [WKT](https://en.wikipedia.org/wiki/Well-known_text_representation_of_geometry) format.
| ## Parameters Use the following parameters to configure the HeatMap gem. | Parameter | Description | | --------------------- | -------------------------------------------- | | Longitude Column Name | Input column that contains longitude values. | | Latitude Column Name | Input column that contains latitude values. | ### Advanced The following table describes the advanced settings for this gem. | Parameter | Description | | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Heat Column Name | Specifies a numeric column to determine the heat intensity at each point.
If not set, each point contributes equally. Optional. | | Decay Function | Defines how heat intensity decreases with distance from the center point.
  • `Constant` applies equal weight to all hexes
  • `Linear` reduces weight proportionally with distance
  • `Exponential` halves the weight at each step away
| | Resolution | Sets the size of each hexagon using the H3 indexing system.
Lower resolutions result in larger hexes, while higher values create finer grids. | | Grid Distance | Specifies how many hexagon steps away from the center should receive heat.
A value of 1 includes immediate neighbors, while higher values expand the influence area. | # FindNearest Source: https://docs.prophecy.ai/data-analysis/gems/spatial/nearest-point Identify the shortest distance between spatial objects This gem runs in . ## Overview Find the closest spatial point(s) between two datasets based on geographic distance. This gem compares each point in the first dataset to all points in the second dataset and returns the nearest matches. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `ProphecyDatabricksSqlSpatial` version 0.0.3 or higher to your project. ## Input and Output | Port | Description | | ------- | ----------------------------------------------------------------------------------------- | | **in0** | Input dataset containing the source points for which you want to find the nearest points. | | **in1** | Input dataset containing the target points to compare against. | | **out** | Output dataset with the nearest point(s) from **in1** for each point in **in0**. | The output schema of **out** contains: * All columns from **in0**. * All columns from **in1**. * `rank_number`: The rank of each match, where `1` is the closest point. * `distance`: The calculated distance between the source and target points in the unit of measurement that you specify. * `cardinal_direction`: The compass direction from the source point to the target point (e.g., `NW` for northwest). Geographic points must be in Well-known Text ([WKT](https://en.wikipedia.org/wiki/Well-known_text_representation_of_geometry)) format. Use the [CreatePoint](/data-analysis/gems/spatial/create-point) gem to convert longitude and latitude coordinates to WKT format. ## Parameters Configure the FindNearest gem using the following parameters. ### Spatial Object Fields | Parameter | Description | | ---------------------- | ------------------------------------------------------------------------- | | Source Centroid Type | Type of geospatial object in **in0**. Currently, only Point is supported. | | Source Centroid Column | Column in **in0** that contains the source spatial points. | | Target Centroid Type | Type of geospatial object in **in1**. Currently, only Point is supported. | | Target Centroid Column | Column in **in1** that contains the target spatial points. | ### Select Output Options | Parameter | Description | | -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | | How many nearest points to find? | Number of closest points to return from **in1** for each point in **in0**. | | Maximum distance | Option to limit the search to target points within this distance (specify units). When the maximum distance is `0`, no maximum is enforced. | | Ignore 0 distance matches | Whether to exclude points that have exactly the same coordinates as the source point. | ## Example: Customer service centers Assume you have two tables: * `customer_locations` that contains the location of each customer.
| `customer_id` | `customer_point` | | ------------- | ------------------------ | | C001 | POINT(-122.4194 37.7749) | | C002 | POINT(-74.0060 40.7128) |
* `service_centers` that contains the location of each service center.
| `center_id` | `center_point` | | ----------- | ------------------------ | | S100 | POINT(-122.4192 37.7793) | | S200 | POINT(-73.9352 40.7306) | | S300 | POINT(-118.2437 34.0522) |
You can use the FindNearest gem to find the nearest service centers to each customer. 1. Add the FindNearest gem to your pipeline canvas. 2. Connect the `customer_locations` table to the FindNearest `in0` input port. 3. Connect the `service_centers` table to the FindNearest `in1` input port. 4. Open the FindNearest gem configuration. 5. For **Source Centroid Type**, select Point. 6. For **Source Centroid Column**, select the `customer_point` column. 7. For **Target Centroid Type**, select Point. 8. For **Target Centroid Column**, select the `center_point` column. For this example, let's find the **two** nearest service centers in a 1000 km radius. 1. Type `2` in the **How many nearest points to find?** field. 2. Type `1000` and choose **Kilometers** in the **Maximum distance** field. 3. Lastly, save and run the gem. ### Result The output contains the two closest service centers to each customer. Ranks begin at `1`, with `1` being the closest point. Note that customer `C002` only has one service center within 1000 km from their location. | `customer_id` | `customer_point` | `center_id` | `center_point` | `rank_number` | `distanceKilometers` | `cardinal_direction` | | ------------- | ------------------------ | ----------- | ------------------------ | ------------- | -------------------- | -------------------- | | C001 | POINT(-122.4194 37.7749) | S100 | POINT(-122.4192 37.7793) | 1 | 0.48957333464416436 | N | | C001 | POINT(-122.4194 37.7749) | S300 | POINT(-118.2437 34.0522) | 2 | 559.1205770615533 | SE | | C002 | POINT(-74.0060 40.7128) | S200 | POINT(-73.9352 40.7306) | 1 | 6.286267237667312 | E | # PolyBuild Source: https://docs.prophecy.ai/data-analysis/gems/spatial/polybuild Create a polygon or polyline from a set of coordinates This gem runs in . ## Overview Build spatial shapes from coordinate data by grouping and ordering points into either polygons (closed shapes) or polylines (open lines). Use this gem to convert raw latitude/longitude values into structured spatial geometries for mapping, analysis, or downstream geospatial operations. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `ProphecyDatabricksSqlSpatial` version 0.0.3 or higher to your project. ## Input and Output The PolyBuild gem accepts the following input and output. | Port | Description | | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **in0** | Input dataset containing pairs of columns with latitude and longitude coordinates, along with fields for grouping and ordering coordinates. | | **out** | Output dataset with one row per group, each containing a generated polygon or polyline in [WKT](https://en.wikipedia.org/wiki/Well-known_text_representation_of_geometry) format. | ## Parameters Configure the PolyBuild gem using the following parameters. | **Parameter** | **Description** | | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Build Method | Select the type of spatial geometry to create from your coordinate data:
  • Select **Sequence Polygon** for closed areas, such as boundaries or zones
  • Select **Sequence Polyline** for open paths, such as routes or trails
| | Longitude Column Name | Column containing longitude values. Must be in decimal degrees (e.g., `-122.4194`). | | Latitude Column Name | Column containing latitude values. Must be in decimal degrees (e.g., `37.7749`). | | Group Field | Column used to divide coordinates into distinct shapes.
Example: Use the `state_name` column to identify polygons for each state border.
*Only one column is supported for grouping.* | | Sequence Field | Column that defines the drawing order of points within each group.
This column can be any sortable type (e.g., integers, timestamps, strings).
The points will be connected in ascending order. | * Sequence Polygon: Uses the `POLYGON()` format to create a closed shape. * Sequence Polyline: Uses the `LINESTRING()` format to create an open path. ## Example Assume you have the following `routes` table for public transportation routes.
| route\_id | stop\_schedule | latitude | longitude | | ---------------- | -------------------- | -------- | --------- | | bus\_21\_morning | 2025-07-16T08:00:00Z | 37.7749 | -122.4194 | | bus\_21\_morning | 2025-07-16T08:10:00Z | 37.7793 | -122.4192 | | bus\_21\_morning | 2025-07-16T08:20:00Z | 37.7796 | -122.4148 | | tram\_5\_evening | 2025-07-16T18:00:00Z | 34.0522 | -118.2437 | | tram\_5\_evening | 2025-07-16T18:10:00Z | 34.0565 | -118.2470 | | tram\_5\_evening | 2025-07-16T18:20:00Z | 34.0580 | -118.2417 |
To transform each route into a polyline geometry: 1. Add the PolyBuild gem to your pipeline canvas. 2. Connect the `routes` table to the PolyBuild input port. 3. Open the PolyBuild gem configuration. 4. For **Build Method**, select **Sequence Polyline**. 5. For **Longitude Column Name**, select the `longitude` column. 6. For **Latitude Column Name**, select the `latitude` column. 7. For **Group Field**, select the `route_id` column. 8. For **Sequence Field**, select the `stop_schedule` column. 9. Save and run the gem. ### Result The PolyBuild gem outputs a table including a polyline for each route.
| route\_id | wkt | | ---------------- | ---------------------------------------------------------------------- | | bus\_21\_morning | `LINESTRING (-122.4194 37.7749, -122.4192 37.7793, -122.4148 37.7796)` | | tram\_5\_evening | `LINESTRING (-118.2437 34.0522, -118.2470 34.0565, -118.2417 34.0580)` |
# Simplify Source: https://docs.prophecy.ai/data-analysis/gems/spatial/simplify Decrease the number of nodes that make up a polygon or polyline This gem runs in . ## Overview The Simplify gem reduces the number of vertices in polygons and polylines while preserving their overall shape. This is useful for reducing file sizes, improving rendering performance, and simplifying complex geometries for analysis or visualization. The gem uses the [Ramer-Douglas-Peucker algorithm](https://en.wikipedia.org/wiki/Ramer%E2%80%93Douglas%E2%80%93Peucker_algorithm), which removes vertices based on their perpendicular distance from line segments. You can control the level of simplification by adjusting the distance threshold. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `ProphecyDatabricksSqlSpatial` version 0.0.4 or higher to your project. ## Input and Output The Simplify gem accepts the following input and output. | Port | Description | | ------- | ------------------------------------------------------------------------------ | | **in0** | Source dataset containing a column with lines or polygons. | | **out** | Output dataset containing a new column with the transformed lines or polygons. | Input geometries must be in [WKT](https://en.wikipedia.org/wiki/Well-known_text_representation_of_geometry) format. Use the [PolyBuild](/data-analysis/gems/spatial/polybuild) gem to convert longitude and latitude coordinates into polygons or polylines in WKT format. ## Parameters Configure the Simplify gem using the following parameters. | Parameter | Description | | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | Geometry column | Column containing the WKT geometries to simplify. | | Tolerance | Distance threshold for vertex removal.
Vertices closer than this distance to the line segment connecting their neighbors will be removed. | | Units | Unit of measurement for the threshold in miles or kilometers. | # SpatialMatch Source: https://docs.prophecy.ai/data-analysis/gems/spatial/spatial-match Find relationships between geographic features This gem runs in . ## Overview Use the SpatialMatch gem to find relationships between geometries from two different datasets. Common use cases include: * Finding which stores are located within specific delivery zones * Identifying roads that intersect with flood zones * Matching customer locations to their nearest service areas The gem uses spatial joins to compare geometries and returns only the pairs that have the spatial relationship you specify, such as shapes overlapping or shapes touching. It works with points, lines, and polygons in Well-Known Text ([WKT](https://en.wikipedia.org/wiki/Well-known_text_representation_of_geometry)) format. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `ProphecyDatabricksSqlSpatial` version 0.0.3 or higher to your project. ## Input and Output The SpatialMatch gem accepts the following inputs and output. | Port | Description | | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **in0** | Dataset containing geometries (points, lines, or polygons) in WKT format. This is the "left" dataset in the spatial join operation. | | **in1** | Dataset containing geometries (points, lines, or polygons) in WKT format. This is the "right" dataset in the spatial join operation. | | **out** | Output dataset containing **matched pairs** of geometries along with all additional columns from both input datasets. Each row represents a source geometry and target geometry that satisfy the selected spatial relationship. Unmatched geometries are excluded from the output.

The output includes the following columns:
  • The source geometry column
  • All other `in0` columns
  • The target geometry column prefixed with `target_`
  • All other `in1` columns prefixed with `target_`
| You can use the same source for both in0 and in1 if you want to match geometries from the same dataset (self-join). Use the following gems to create correctly formatted geometries in a dataset: * [CreatePoint](/data-analysis/gems/spatial/create-point) gem for points * [PolyBuild](/data-analysis/gems/spatial/polybuild) gem for lines and polygons ## Parameters Configure the SpatialMatch gem using the following parameters. | Parameter | Description | | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | | Source Column | Select the column from `in0` that contains one set of geometric data. | | Target Column | Select the column from `in1` that contains another set of geometric data. | | Select Match Type | Choose the spatial relationship that determines when a source geometry matches a target geometry.
Learn more in [Match types](#match-types). | ### Match types Review the following to understand the criteria to satisfy different match types. The SpatialMatch gem returns a row for each match condition that is met. | Match type | Description | | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Source Intersects Target | Condition is met if the source and target geometries share any portion of space. This is the most general spatial relationship. Any overlap, touching, or containment satisfies intersection. | | Source Contains Target | Condition is met if the source geometry completely contains the target geometry. The target geometry must be entirely within the source geometry's interior and boundary. | | Source Within Target | Condition is met if the source geometry is completely contained within the target geometry. This is the inverse of the **Contains** relationship. | | Source Touches Target | Condition is met if the source and target geometries have at least one point in common, but their interiors do not intersect. | | Source Touches or Intersects Target | Condition is met if the source and target geometries either touch (share boundary points) or intersect (share any portion of space). This combines the touch and intersect relationships. | | Source Envelope Overlaps Target Envelope | Condition is met if the minimum bounding rectangles (envelopes) of the source and target geometries overlap. This is a less precise check than a standard intersection. | #### Match types diagram The following diagram includes visualizations for each match type. Match types diagram ## Example: Find stores within delivery zones Assume you have two datasets: * `store_locations` contains store locations as points.
| store\_id | store\_name | store\_location | store\_type | | --------- | ------------------------ | ----------------------- | ----------- | | 1 | Downtown Electronics | POINT(-74.0059 40.7128) | electronics | | 2 | Midtown Cafe | POINT(-73.9857 40.7489) | restaurant | | 3 | Brooklyn Bookstore | POINT(-73.9442 40.6782) | bookstore | | 4 | Queens Pharmacy | POINT(-73.7949 40.7282) | pharmacy | | 5 | Upper East Side Boutique | POINT(-73.9626 40.7831) | clothing |
* `delivery_zones` contains delivery zones as polygons.
| zone\_id | zone\_name | zone\_polygon | delivery\_fee | | ---------------- | ---------------- | --------------------------------------------------------------------------------------------------- | ------------- | | MANHATTAN\_SOUTH | Lower Manhattan | POLYGON((-74.0200 40.7000, -74.0200 40.7300, -73.9800 40.7300, -73.9800 40.7000, -74.0200 40.7000)) | 5.99 | | MANHATTAN\_NORTH | Upper Manhattan | POLYGON((-73.9800 40.7700, -73.9800 40.8000, -73.9400 40.8000, -73.9400 40.7700, -73.9800 40.7700)) | 7.99 | | BROOKLYN\_WEST | Western Brooklyn | POLYGON((-74.0000 40.6500, -74.0000 40.7000, -73.9200 40.7000, -73.9200 40.6500, -74.0000 40.6500)) | 6.99 | | QUEENS\_CENTRAL | Central Queens | POLYGON((-73.8500 40.7000, -73.8500 40.7500, -73.7500 40.7500, -73.7500 40.7000, -73.8500 40.7000)) | 8.99 |
To find which delivery zones correspond to each store: 1. Add a SpatialMatch gem to your pipeline canvas. 2. Attach `store_locations` to the **in0** port of the gem. 3. Attach `delivery_zones` to the **in1** port of the gem. 4. Open the gem configuration interface. 5. For the **Source** field, select the `store_location` column from the `store_locations` table. 6. For the **Target** field, select the `zone_polygon` column from the `delivery_zones` table. 7. Under **Select Match Type**, select **Source Within Target**. 8. Save and run the gem. ### Result The SpatialMatch gem will return only the pairs of geometries that satisfy the selected match type. As a result: * Downtown Electronics matches the `MANHATTAN_SOUTH` zone. * Queens Pharmacy matches the `QUEENS_CENTRAL` zone. * Upper East Side Boutique matches the `MANHATTAN_NORTH` zone. * Midtown Cafe does not match any zone. This means that it is not in any delivery zone. # Aggregate gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/transform/aggregate Group and summarize your data using GROUP BY and aggregation functions This gem runs in . ## Overview By default, the Aggregate gem provides a guided interface for defining grouping columns, aggregation functions, and aggregate filters. If you need more complex expressions, you can switch to Advanced mode to edit the underlying SQL configuration directly. The Aggregate gem transforms many input rows into fewer output rows. It groups records that share the same values and calculates summary statistics for each group. Common use cases include: * Counting orders per customer. * Calculating total sales by region. * Finding average transaction values. * Summarizing records by category or date. * Creating grouped metrics for dashboards and downstream analysis. The Aggregate gem works similarly to a SQL `GROUP BY` statement. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Modes The Aggregate gem supports two editing modes: | Mode | Description | | ------------------------- | ----------------------------------------------------------------------------------------------------------------- | | **Simple mode** (default) | Build aggregations using guided controls for grouping, aggregation functions, and aggregate filters. | | **Advanced mode** | Edit the full Aggregate configuration directly for complex expressions that cannot be represented in Simple mode. | ## Simple mode Simple mode organizes the Aggregate gem into three collapsible sections. | Section | Description | | --------------------- | ---------------------------------------------------------------------------------------------------------- | | **Grouping** | Choose one or more columns to group by. Optionally, rename the output column for each grouping expression. | | **Aggregations** | Select an aggregation function, choose the input column, and provide an output column name. | | **Aggregate Filters** | Filter grouped results after aggregation using a HAVING expression. | ## Supported aggregation functions | Function | Description | | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | count | Counts rows. | | count distinct | Counts distinct values. | | sum | Returns the total. | | avg | Returns the average. | | min | Returns the minimum value. | | max | Returns the maximum value. | | any\_value | Returns an arbitrary value from the group. Use `any_value` when every row in a group has the same value, or when any representative value is acceptable. | | stddev | Returns the standard deviation. | | variance | Returns the variance. | | collect\_list | Returns an array containing all values. | | collect\_set | Returns an array containing unique values. | ## Complex expressions Simple mode supports standard grouping columns and aggregation expressions such as `sum(order_amount)` or `avg(age)`. If an existing Aggregate gem contains more advanced expressions, Simple mode displays a message indicating that the configuration is too complex to edit in Simple mode. For example, Simple mode does note support calculated grouping expressions such as `sum(a) + sum(b)`. In these cases, you have two options: * **Switch to Advanced Mode** lets you edit the existing configuration in Advanced Mode. * **Reset and use Default Mode** clears the grouping and aggregation configurations so that you can start over. **Reset and use Default Mode** permanently removes the existing grouping and aggregation configuration. Aggregate Filters (HAVING conditions) are preserved. ## How the Aggregate gem works The Aggregate gem processes data in three main steps: 1. Group rows based on the selected **Group By** columns. 2. Calculate aggregation expressions for each group. 3. Optionally filter grouped results using **Having Conditions**. ### Example grouping | Department | Average salary | Employee count | | ---------- | -------------: | -------------: | | HR | 75 | 2 | | IT | 92.5 | 2 | | Sales | 60 | 1 | After applying `employee_count >= 2`, `Sales` is removed. ```mermaid theme={null} flowchart LR A[Input rows] --> B[Group by Department] B --> C[Calculate AVG Salary and COUNT Employees] C --> D[Intermediate grouped results] D --> E{employee_count >= 2?} E -- Yes --> F[Keep group] E -- No --> G[Remove group] ``` Unlike a Filter gem, **Having Conditions** are evaluated after aggregation is complete. For example: * A Filter gem can remove individual rows before grouping. * A Having condition can remove grouped results after calculations are complete. ## Common use cases | Goal | Group By | Expression | | ------------------------------------------- | ------------- | ---------------------------- | | Count orders per customer | `customer_id` | `count(order_id)` | | Calculate total revenue by region | `region` | `sum(order_amount)` | | Find average order value by sales rep | `sales_rep` | `avg(order_amount)` | | Get the latest transaction date per account | `account_id` | `max(transaction_date)` | | Count distinct products sold by store | `store_id` | `count_distinct(product_id)` | ## Example Suppose you have a dataset of orders that includes heathcare users with the following columns: | Column name | Description | | ------------------------- | -------------------------------------------------------------------------- | | `persona` | Customer or employee segment used for demographic and behavioral analysis. | | `age` | Employee age in years. | | `annual_household_income` | Total annual household income. | | `dependents_on_plan` | Number of dependents enrolled in the employee's healthcare plan. | You can use the Aggregate gem to summarize users by `persona` and calculate average demographic metrics for each segment. ### Example configuration | Grouping | Aggregation | Output column | | --------- | ------------------------------ | ------------------------------- | | `persona` | `avg(annual_household_income)` | `average_household_income` | | | `avg(age)` | `average_age` | | | `avg(dependents_on_plan)` | `avg_number_dependents_on_plan` | | | `any_value(persona)` | `persona` | ### Result This configuration: * Groups all rows by `persona`. * Calculates the average household income, age, and number of dependents for each persona. * Returns the `persona` value for each group using `any_value`. Because `persona` is also the grouping column, every row within a group has the same `persona` value. In cases like this, `any_value` returns that shared value and includes it in the output. ## Work with grouping and aggregation rows Both the **Grouping** and **Aggregations** sections support: * Searching for existing rows. * Adding rows with **+**. * Deleting rows. * Dragging rows to reorder them. Hovering over a row highlights the corresponding source column in the input preview and the generated column in the output preview. Clicking a row keeps that highlight active until you select another row. ## Common issues ### Unexpected duplicate groups Check for: * Leading or trailing spaces in grouped columns. * Differences in letter casing such as `US` versus `us`. * Null values creating separate groups. You may need to clean or standardize values before aggregation. ### Incorrect counts If counts appear too high: * Verify whether duplicate rows exist before aggregation. * Confirm whether you should use `count()` or `count_distinct()`. ### Aggregate Filters/Having Conditions does not work as expected Aggregate Filters (Simple mode) and Having Conditions (Advanced mode) filter grouped results after aggregation. If you need to filter raw rows before grouping, use a Filter gem earlier in the pipeline. ## Similar tools and concepts | Tool or Platform | Similar Concept | | ---------------- | ------------------------------------- | | SQL | `GROUP BY` with aggregation functions | | Alteryx | Summarize tool | | Pandas | `groupby()` with aggregate functions | | PySpark | `groupBy().agg()` | | Excel | PivotTables and grouped summaries | # CountRecords Source: https://docs.prophecy.ai/data-analysis/gems/transform/count-records Returns one integer that represents the count of records in the input dataset This gem runs in . ## Overview The CountRecords gem allows you to count the number of rows in a dataset in different ways. You can count all rows, count non-null values in selected columns, or count distinct non-null values in selected columns. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Input and Output The CountRecords gem accepts the following input and output. | Port | Description | | ------- | ------------------------------------------------------------------------------------------ | | **in0** | Input dataset with the columns to count. | | **out** | Output dataset with the resulting count(s). Output has one row with the selected count(s). | ## Parameters Configure the CountRecords gem using the following parameters. | Parameter | Description | | ----------------------- | ----------------------------------------------------------------------------------------- | | Count option | Choose how the data should be counted. See [Count options](#count-options) below. | | Select columns to count | One or more columns to count. Required for counting non-null records or distinct records. | ### Count options Choose one of the following strategies for counting records. | Strategy | Description | | -------------------------------------------- | ----------------------------------------------------------------------------- | | Count number of total records | Returns the total number of rows in the input dataset, including null values. | | Count non-null records in selected column(s) | Returns the number of non-null rows for each selected column. | | Count distinct records in selected column(s) | Returns the number of distinct, non-null values for each selected column. | ## Example Given a table of patient visits:
| PatientID | VisitDate | Department | Diagnosis | | --------- | ---------- | ---------- | --------- | | 1 | 2024-01-01 | Cardiology | Flu | | 2 | 2024-01-02 | Oncology | Cancer | | 3 | 2024-01-03 | Cardiology | Flu | | 4 | 2024-01-04 | NULL | Cold |
If you choose: * **Count distinct records** on `Department`: the result will be `2` (Cardiology, Oncology). * **Count non-null records** on `Department`: the result will be `3`. * **Count total number of records**: the result will be `4`. # DynamicSelect gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/transform/dynamic-select Dynamically filter columns of your dataset based on a set of conditions This gem runs in . ## Overview Use the DynamicSelect gem to dynamically filter columns of your dataset based on a set of conditions to avoid hard-coding your choice of columns. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Parameters | Parameter | Description | | ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Configuration | Whether to filter the columns visually or with code.
  • Select field types: Use checkboxes to select column types to keep in the dataset, such as string, decimal, or date.
  • Select via expression: Create an expression that limits the type of columns to keep in the dataset.
| ## Example Let's say you have the following table with weather prediction data.
| DatePrediction - `Date` | TemperatureCelsius - `Integer` | HumidityPercent - `Integer` | WindSpeed - `Float` | Condition - `String` | | ----------------------- | ------------------------------ | --------------------------- | ------------------- | -------------------- | | 2025-03-01 | 15 | 65 | 10.0 | Sunny | | 2025-03-02 | 17 | 70 | 12.2 | Cloudy | | 2025-03-03 | 16 | 68 | 11.0 | Rainy | | 2025-03-04 | 14 | 72 | 9.8 | Sunny |
### Remove columns using field type Assume you would like to remove irrelevant float and string columns from your dataset. You can do so with the **Select field types** method by selecting all field types to maintain, except for float and string. ### Remove columns with an expression Using the same example, you can accomplish the same task with the **Select via expression** method by inputting the the expression `column_type NOT IN ('Float', 'String')`. Be aware that column types are case sensitive. Use the same format shown in the input table schemas in the gem configuration. ### Result
| DatePrediction - `Date` | TemperatureCelsius - `Integer` | HumidityPercent - `Integer` | | ----------------------- | ------------------------------ | --------------------------- | | 2025-03-01 | 15 | 65 | | 2025-03-02 | 17 | 70 | | 2025-03-03 | 16 | 68 | | 2025-03-04 | 14 | 72 |
# DataEncoderDecoder Source: https://docs.prophecy.ai/data-analysis/gems/transform/encoder-decoder Encode and decode data using different techniques This gem runs in . ## Overview The DataEncoderDecoder gem allows you to encode or decode data in selected columns using a variety of standard techniques, including Base64, Hex, and AES encryption. You can transform values in-place or create new output columns with a prefix or suffix. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Input and Output The DataEncoderDecoder gem accepts the following input and output. | Port | Description | | ------- | ------------------------------------------------------------------------------------------------------------------------------ | | **in0** | Input dataset containing one or more string or binary columns to encode or decode. | | **out** | Output dataset with transformed values. Output columns depend on the [transformed column option](#transformed-column-options). | ## Parameters Configure the DataEncoderDecoder gem using the following parameters. | Parameter | Description | | ------------------------------- | ------------------------------------------------------------------------------------------------------- | | Select columns to encode/decode | One or more columns to apply the transformation to. | | Select encode/decode option | The encoding or decoding method to apply. See [Encode/Decode methods](#encodedecode-methods). | | Transformed column options | Choose how the output should be written. See [Transformed column options](#transformed-column-options). | ### Encode/Decode methods Choose from the following methods: #### `base64` Encodes the selected column(s) using Base64. #### `unbase64` Decodes Base64-encoded column values. #### `hex` Encodes the selected column(s) into hexadecimal format. #### `unhex` Decodes hexadecimal-encoded values. #### `encode` Encodes the string column(s) using a specified character set. * **Charset**: Character set to use, such as `UTF-8`. #### `decode` Decodes the string column(s) using a specified character set. * **Charset**: Character set to use, such as `UTF-8`. #### `aes_encrypt` Encrypts the selected column(s) using AES encryption. * **Secret scope**: The name of the Databricks secret scope. * **Secret key**: The key name within the scope that stores the encryption key. * **Mode**: AES encryption mode to use. * `GCM` * `CBC` * `EBC` * **(Optional) AAD scope and key**: For GCM mode, you can specify Databricks scope and key for the AAD. * **(Optional) Initialization vector scope and key**: For CBC mode, specify a Databricks secret scope and key for the IV. ### Transformed column options * **Substitute the new columns in place**: Replaces the original column(s) with the transformed values. * **Add new columns with a prefix/suffix attached**: Adds a new column for each transformed input column, appending a prefix or suffix to the name. ## Example Assume you have the following dataset:
| `ID` | `Message` | | ---- | ------------ | | 1 | Hello world! | | 2 | Prophecy |
Using the `base64` method, adding new columns with the suffix `_encoded`, the output would be:
| `ID` | `Message` | `Message_encoded` | | ---- | ------------ | ----------------- | | 1 | Hello world! | SGVsbG8gd29ybGQh | | 2 | Prophecy | UHJvcGhlY3k= |
# FuzzyMatch gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/transform/fuzzy-match Match records that are not exactly identical This gem runs in . ## Overview Use the FuzzyMatch gem to identify non-identical duplicates in your data. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Prerequisites * Add `prophecy_basics` package version 1.0.0 or higher to your project. ## Input and Output | Table | Description | | ------- | ----------------------------------------------------------------------------------------------------- | | **in0** | Includes the table on which duplicates will be checked.
Note: FuzzyMatch only allows one input. | | **out** | Generates one record per fuzzy match. | ## Parameters ### Configuration | Parameter | Description | | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Merge/Purge Mode | Records are either compared from a single source (Purge) or across multiple sources (Merge).
Merge mode assumes that multiple sources exist in the same table **in0**. | | Source ID Field | Unique identifier for each source when using **Merge** mode.
This is necessary because the different sources exist in the same table **in0**. | | Record ID Field | Unique identifier for each record. | | Match threshold percentage | If the match score is less than the threshold, the record does not qualify as a match. | | Include similarity score | Checkbox to enable for an additional output column that includes the similarity score. | ### Match Fields | Parameter | Description | | -------------- | --------------------------------------------------------- | | Field name | Name of the column that you want to check for duplicates. | | Match function | The method that generates the similarity score. | ## Example One common use case for the FuzzyMatch gem is to match similarly spelled names. Here's a table with two entries for `Alex Taylor`, whose phone number was updated.
| `id` | `email` | `phone` | `first_name` | `last_name` | `date_added` | | ---- | --------------------- | ------------ | ------------ | ----------- | ------------ | | 1 | `alex.t@example.com` | 123-456-7890 | Alex | Taylor | 2023-01-01 | | 2 | `alex.t@example.com` | 123-456-9542 | Alex | Ttaylor | 2023-07-01 | | 3 | `sam.p@example.com` | 987-654-3210 | Sam | Patel | 2024-03-15 | | 4 | `casey.l@example.com` | 555-111-2222 | Casey | Lee | 2024-05-01 |
You can use the FuzzyMatch gem to find the closely spelled name. In the gem configuration: 1. Set the Merge/Purge Mode to **Purge mode**. 2. For the Record ID, use the **id** column. 3. Keep the threshold at `80` percent. 4. Enable the **Include similarity score column** checkbox. 5. In the Match Fields tab, add a match field for the **last\_name** column. 6. Set the Match Function to **Name**. 7. Save and run the gem. ### Result The output includes the Record IDs of the records with fuzzy matches above the defined threshold.
| `id` | `id2` | `similarityScore` | | ---- | ----- | ------------------ | | 1 | 2 | 0.9111111111111111 |
Depending on your SQL provider, you might see different similarity scores based on the algorithm that runs under the hood. To view the names per record, [join](/data-analysis/gems/join-split/join) the FuzzyMatch output with the original dataset. # Pivot Source: https://docs.prophecy.ai/data-analysis/gems/transform/pivot Convert your table from long to wide format This gem runs in . Use the Pivot gem to convert your table from long format to wide format. You select a column whose unique values become new columns in the output, making it easier to summarize and compare categories side by side. Each unique value that you select from the pivot column creates a new column in the result. ## Input and output The Pivot gem uses the following input and output ports. | Port | Description | | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **in0** | The source table containing data in long format that needs to be transformed to wide format. | | **out** | The output table in wide format containing:
β€’ All original grouping columns preserved as row identifiers
β€’ New columns created from unique values in the pivot column and defined aggregations
The output will typically have fewer rows but more columns than the input. | ## Parameters Configure the Pivot gem using the parameters for your SQL warehouse. ### Grouping Select one or more columns to group your data before applying the aggregation. You can also change the grouping column name in the output. For example, to analyze sales data by region and product category, select both `Region` and `Product_Category` as grouping columns. ### Pivot Choose a pivot column that contains the unique values that will become new columns in the output table. After you select the column, click **Fetch unique values** to retrieve all distinct values from that column. Select the values that you want to transform into new columns. ### Aggregation Choose a column containing the data values that will populate the cells of the new pivot columns. The values are aggregated and placed at the intersection of each group and pivot column combination. Choose an aggregation method such as `AVG`, `SUM`, `COUNT`, `MIN`, or `MAX`. You cannot define a column alias for the aggregation in Snowflake. ### Grouping Select one or more columns to group your data before applying the aggregation. You can also change the grouping column name in the output. For example, to analyze sales data by region and product category, select both `Region` and `Product_Category` as grouping columns. ### Pivot Choose a pivot column that contains the unique values that will become new columns in the output table. After you select the column, click **Fetch unique values** to retrieve all distinct values from that column. Select the values that you want to transform into new columns. ### Value for New column Choose a column containing the data values that will populate the cells of the new pivot columns. The values are aggregated and placed at the intersection of each group and pivot column combination. ### Aggregation function Choose an aggregation method such as `AVG`, `SUM`, `COUNT`, `MIN`, or `MAX`. You can define an optional **Column Alias** that is appended to the new column names to make the output columns more descriptive. ### Grouping Select one or more columns to group your data before applying the aggregation. You can also change the grouping column name in the output. For example, to analyze sales data by region and product category, select both `Region` and `Product_Category` as grouping columns. ### Pivot Choose a pivot column that contains the unique values that will become new columns in the output table. After you select the column, click **Fetch unique values** to retrieve all distinct values from that column. Select the values that you want to transform into new columns. ### Value for New column Choose a column containing the data values that will populate the cells of the new pivot columns. The values are aggregated and placed at the intersection of each group and pivot column combination. ### Aggregation function Choose an aggregation method such as `AVG`, `SUM`, `COUNT`, `MIN`, or `MAX`. You can define an optional **Column Alias** that is appended to the new column names to make the output columns more descriptive. ## Example: Energy consumption Imagine you have a dataset of energy consumption with `Date`, `City`, and `kWh_Used` for different buildings, and you want to see the average `kWh_Used` per city for specific dates, with dates as columns. You start with the following data:
| Date | Building\_ID | Building\_Type | City | Energy\_Source | kWh\_Used | Cost | Peak\_Demand\_kWh | | ---------- | ------------ | -------------- | ----------- | -------------- | --------- | ----- | ----------------- | | 2025-07-01 | B001 | Office | New York | Solar | 150 | 30.00 | 10 | | 2025-07-01 | B002 | Residential | Los Angeles | Grid | 200 | 40.00 | 12 | | 2025-07-01 | B003 | Retail | New York | Grid | 180 | 36.00 | 11 | | 2025-07-02 | B001 | Office | New York | Solar | 160 | 32.00 | 10.5 | | 2025-07-02 | B002 | Residential | Los Angeles | Grid | 210 | 42.00 | 12.5 | | 2025-07-03 | B004 | Commercial | Chicago | Grid | 250 | 50.00 | 15 | | 2025-07-03 | B005 | Industrial | Houston | Solar | 300 | 60.00 | 18 | | 2025-07-04 | B001 | Office | New York | Solar | 155 | 31.00 | 10.2 | | 2025-07-04 | B003 | Retail | New York | Grid | 185 | 37.00 | 11.5 | | 2025-07-05 | B002 | Residential | Los Angeles | Grid | 205 | 41.00 | 12.3 |
In the Pivot gem configuration: 1. For **Grouping**, select the `City` column. 2. For **Pivot**, select the `Date` column. 3. Click **Fetch unique values**, and select `'2025-07-01'`, `'2025-07-02'`, `'2025-07-03'`, `'2025-07-04'`, and `'2025-07-05'`. 4. Select `kWh_Used` as the column containing the values for the new columns. 5. Select `AVG` as the aggregation method. ### Result The output table shows the average `kWh_Used` for each city on the specified dates, with the dates as new columns. Some cells are null because there is no corresponding data for certain combinations of `City` and `Date` in the input table. The Pivot gem turns the selected `Date` values into columns and populates them with aggregated `kWh_Used` values for each `City`. For example, Chicago has no data for `2025-07-01`, so its value for that date is null. A null value therefore indicates that the input dataset contains no data for that city and date combination. # RunningTotal Source: https://docs.prophecy.ai/data-analysis/gems/transform/running-total Calculate running totals for selected numeric columns. This gem runs in . Use the RunningTotal gem to calculate running totals for selected numeric columns. You can also configure optional partitioning and row ordering for the calculation. This gem is useful when you want to control how running total calculations are grouped and ordered across your data. ## Prerequisites * Add `prophecy_basics` package version 1.0.11 or higher to your project. ## Parameters | Parameter | Description | | ------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | | Columns for running total | Select the columns to include in the running total calculation. | | Partition by (optional) | Optionally select fields to partition the calculation. | | Order rows for calculation (optional) | Configure how rows are ordered for the calculation. This section includes **Order By Columns** and **Sort strategy**. | | Output column prefix (optional) | Enter a prefix for the output columns. | ## Order rows for calculation Use **Order By Columns** to select the columns used to order rows for the calculation. Use **Sort strategy** to choose one of the following options: * ascending nulls first * ascending nulls last * descending nulls first * descending nulls last ## Example Let's say you have the following dataset.
| Record | Value | | ------ | ----- | | A | 10 | | B | 12 | | C | 4 |
If you select **Value** in **Columns for running total**, the gem calculates a cumulative total across the rows. ### Result
| Record | Value | RunningTotal\_Value | | ------ | ----- | ------------------- | | A | 10 | 10 | | B | 12 | 22 | | C | 4 | 26 |
# Unpivot gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/transform/unpivot Convert your table from wide to long format This gem runs in . ## Overview The Unpivot gem is a data transformation tool designed to reshape datasets from a wide to a long format. This operation consolidates multiple columns containing similar data into two new columns: one representing the original column names, and another holding their corresponding values. This transformation is fundamental for various use cases, particularly for machine learning data preparation. This functionality was previously available through the Transpose gem. While existing pipelines utilizing the Transpose gem will continue to operate, we recommend migrating to the Unpivot gem for new implementations. ## Input and Output The Unpivot gem uses the following input and output ports. | Port | Description | | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **in0** | The source table containing data in wide format, where multiple columns hold similar types of data that need to be consolidated. | | **out** | The output table in long format containing:
β€’ All original key columns (identifiers)
β€’ A new column containing the names of the unpivoted columns. Default column name: `Name`
β€’ A new column with the corresponding values from the unpivoted columns. Default column name: `Value`
The output will be in long format and have more rows than the input. | ## Parameters Configure the Unpivot gem using the following parameters. ### Key Columns Key columns serve as unique identifiers and will be preserved in the output. These columns remain unchanged during the unpivot operation and help maintain the relationship between the original and transformed data. Select all columns that you want to keep as identifiers for each row. > Example: In a test scores dataset with columns `Student_ID`, `Subject`, `Test1`, `Test2`, `Test3`, you would select `Student_ID` and `Subject` as key columns to maintain student and subject identifiers in your transformed data. ### Data Columns Data columns contain the actual data values that you want to transform from wide format (multiple columns) to long format (single column). These columns will be transformed into two new columns: one containing the column names and another containing their values. > Example: In a test scores dataset with columns `Student_ID`, `Subject`, `Test1`, `Test2`, `Test3`, you would select `Test1`, `Test2`, and `Test3` as data columns to transform multiple test scores into a single column structure where each test becomes a separate row. ### Use custom output column names When **Use custom output column names for Name & Value pairs** is enabled, you can override the default column names (`Name` and `Value`) with custom names that better describe your data. > Example: In a test scores dataset with columns `Student_ID`, `Subject`, `Test1`, `Test2`, `Test3`, enable this option and rename **Name** to `Test_Number` and **Value** to `Score` for better semantic meaning in your transformed dataset. ## Example: Sales per quarter Imagine you have sales data for different products, with each quarter's sales (units sold) stored in its own column. This structure is known as wide format. Before modeling seasonal trends or doing time series analysis, it's often helpful to convert this into long format, where each row represents a **single observation**.
| `Product` | `Q1` | `Q2` | `Q3` | `Q4` | | --------- | ---- | ---- | ---- | ---- | | A | 100 | 150 | 130 | 170 | | B | 90 | 120 | 110 | 160 |
To configure a Unpivot gem for this table: 1. For **Key Columns**, select the `Product` column. This allows you to identify quarterly sales per product. 2. For **Data Columns**, select all of the quarter columns (`Q1`, `Q2`, etc.) for your data columns. 3. Select the **Use custom output column names for Name & Value pairs** checkbox. 4. Rename the **Name** column to `Quarter`. 5. Rename the **Value** column to `Units_Sold`. 6. Save and run the gem. ### Result After the transformation: * The quarter names (`Q1`, `Q2`, etc.) will move into a new `Quarter` column. * The corresponding units sold per quarter will be stored in a `Units_Sold` column.
| `Product` | `Quarter` | `Units_Sold` | | --------- | --------- | ------------ | | A | Q1 | 100 | | A | Q2 | 150 | | A | Q3 | 130 | | A | Q4 | 170 | | B | Q1 | 90 | | B | Q2 | 120 | | B | Q3 | 110 | | B | Q4 | 160 |
# WeightedAverage Source: https://docs.prophecy.ai/data-analysis/gems/transform/weighted-average Calculate the weighted average of a numeric field using another numeric field as the weight This gem runs in . ## Overview Use the WeightedAverage gem to calculate the weighted average of a numeric field by using another numeric field as the weight. You can also calculate weighted averages separately for groups of records. The output is an aggregated result set. If no grouping fields are selected, the gem returns a single row containing the weighted average. If grouping fields are selected, it returns one row per group with the grouping fields and the calculated weighted average. ## Prerequisites Add `prophecy_basics` package version 1.0.11 or higher to your project. ## Parameters | Parameter | Description | | -------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Value Field (Numeric) | Select the numeric field that contains the values to average. | | Weight Field (Numeric) | Select the numeric field that contains the weights used in the calculation. When all records have the same weight, the result is equivalent to a standard average. | | Output Field Name | Enter the name of the output field that stores the calculated weighted average. | | Grouping Fields (Optional) | Optionally select one or more fields to calculate the weighted average separately for each group. If you do not select grouping fields, the gem calculates a single weighted average across all input records. | ## How it works The weighted average is calculated as: $$ \frac{\sum (value \times weight)}{\sum weight} $$ Each value is multiplied by its corresponding weight. The results are summed, and that sum is divided by the total weight. If you specify grouping fields, the gem performs this calculation independently for each group. ## Output If you do not select grouping fields, the gem returns a single row and column with the weighted average. If you select grouping fields, the gem returns one row per group with the weighted average. The output includes the selected grouping fields, if any, and the calculated weighted average using the specified output field name. ## Notes * Both the Value column and the Weight column must be numeric. * When weights are equal across all records, the result is the same as a standard average. * Grouping fields are optional and allow per-group calculations. ## Example You want to calculate the weighted average price of items sold, where each price is weighted by `total_sale` value. Input:
| Item | Price | `total_sale` | | ------ | ----- | ------------ | | Shirt | 20 | 2 | | Pants | 50 | 1 | | Jacket | 100 | 3 |
Configuration: * **Value Field**: `Price` * **Weight Field**: `Total_Sale` * **Output Field Name**: `weighted_avg_price` * **Grouping Fields**: *(none)* Calculation: * `(20 Γ— 2 + 50 Γ— 1 + 100 Γ— 3) / (2 + 1 + 3)` * `(40 + 50 + 300) / 6 = 390 / 6 = 65` Output:
| weighted\_avg\_price | | -------------------- | | 65 |
### Example with grouping If you group by `Category`:
| Category | Price | Total\_Sale | | --------- | ----- | ----------- | | Tops | 20 | 2 | | Tops | 50 | 1 | | Outerwear | 100 | 3 |
Configuration: * **Grouping Fields**: `Category` Output:
| Category | `weighted_avg_price` | | --------- | -------------------- | | Tops | 30 | | Outerwear | 100 |
# WindowFunction gem for Data Analysis Source: https://docs.prophecy.ai/data-analysis/gems/transform/window Create moving aggregations and transformation This gem runs in . ## Overview The **WindowFunction** gem lets you perform calculations for individual records across a defined window of rows. This enables transformations such as moving averages, rankings, lead/lag comparisons, and cumulative metrics. Unlike regular aggregations, window functions retain the original row structure while adding new columns that reflect insights based on neighboring rows. You can configure the window using partitions, sorting, and frame definitions to control how each calculation is applied. This page explains how to configure partitions, sort data, and choose frames to achieve the intended behavior. The gem has a corresponding interactive gem example. See [Interactive gem examples](/data-analysis/gems/gems#interactive-gem-examples) to learn how to run sample pipelines for this and other gems. ## Use cases Here are a few common use cases for window functions. | Use case | Example | | ----------------------------------------- | -------------------------------- | | Compute metrics over rolling time periods | Find 7-day average | | Rank or number rows within groups | Rank city population per country | | Compare values from previous or next rows | Find week-over-week change | | Apply cumulative calculations | Get running totals | ## Input and Output The WindowFunction gem uses the following input and output ports. | Port | Description | | ------- | --------------------------------------------------------------------------------------------------------------- | | **in0** | The input dataset containing the columns you want to analyze using window functions. | | **out** | The output dataset, which includes all columns from `in0` and new columns from [window functions](#window-use). | This gem only supports a single input and output port. ## Parameters Explore the following sections to understand how to configure the WindowFunction gem. ### Partition By The **Partition By** tab allows you to define whether you want to partition your data. When you define partitions, windows will apply starting with each group. If no partition is given, one window operates over the whole dataset. You can define multiple partition columns to make your groupings more granular. Columns can be added and kept as is, or you can use an expression to define the column. You can specify one or more columns to partition by. This is useful for calculating things like rankings or rolling metrics within categories, such as per user, region, or product. You can also use expressions to define custom partition logic. > Example: Partitioning by `customer_id` when calculating `row_number()` will restart the row count at 1 for each customer. ### Order By The **Order By** tab controls the order in which rows are processed within each partition. This ordering is essential for functions that depend on sequence, such as calculating rankings, lead or lag time, and rolling averages by date. If no order is specified, the window function output is nondeterministic, since row order may vary between runs. You can sort by column types including numeric, string, date, or timestamp. You can also control whether the sort is ascending or descending, and how null values are treated. > Example: To calculate a rolling average by date, order by a date column in ascending order. You must define an Order By column to use a [range frame](#range-frame) window frame. ### Window Frame The Window Frame defines the subset of rows used to compute each window function result. When no window is specified, the window includes all rows in the partition up to the current row. However, you can customize the frame to include only a specific range of rows. Prophecy supports both [Row Frame](#row-frame) and [Range Frame](#range-frame) options, which are explained below. #### Row Frame The **Row Frame** option defines the window size based on the number of rows before and after the current row. This is useful when you want a fixed number of rows in your window, such as calculating a 3-row moving average, regardless of the values in the ordered column. You configure the frame by setting a **Start** and an **End** boundary. The boundaries can be: * Unbounded Preceding: Includes all rows before the current row. * Unbounded Following: Includes all rows after the current row. * Current Row: Starts or ends at the current row. * Row Number: A specific number of rows before or after the current row. When using row offset, use: * Negative numbers to go backward (e.g., -2 means 2 rows before) * Positive numbers to go forward (e.g., 2 means 2 rows after) * 0 to refer to the current row. > Example: A 2-row moving average might use a frame of `-2` to the current row. This frame includes the current row and the two rows before it in the ordered set. #### Range Frame The Range Frame defines the window using values in the Order By column. This allows you to frame your window based on a value range around the current row, such as a date range or a numerical difference. This is useful when your data has gaps or uneven intervals. You configure the frame by setting a **Start** and an **End** boundary. The boundaries can be: * Unbounded Preceding: Includes all rows before the current row. * Unbounded Following: Includes all rows after the current row. * Current Row: Starts or ends at the current row. * Range Value: A numeric or interval value defining how far before or after the current row to look. Prophecy automatically interprets the unit of the **Range Value** based on the type of the **Order By** column: * If ordering by a date or timestamp, a value of 1 means "1 day forward", and -1 means "1 day backward". * If ordering by a numeric column, 1 means "up to 1 unit above", and -1 means "up to 1 unit below". Because the Range Frame depends on actual data values (not row positions), the number of rows in the window may vary for each record. For example, some window frames may include only one match within 5 days, while others may include ten. ### Window Use In the **Window Use** tab, you can apply window functions that calculate values for each row based on a the defined window frame of surrounding rows. Each function you configure adds a new column to the output dataset. You can apply functions such as: * Ranking functions: `row_number()`, `rank()`, `dense_rank()` and `ntile()`. * Analytical functions: `lead()`, `lag()`, and `cume_dist()`. * Aggregate functions: `min()`, `max()`, and `avg()`. You can add multiple window functions within a single WindowFunction gem. However, all these functions will share the same partitioning and ordering configuration. If you need different partition or order columns for different window calculations, use multiple WindowFunction gems sequentially. For a full list of available functions, visit the [Window functions](https://docs.databricks.com/aws/en/sql/language-manual/sql-ref-window-functions#parameters) page of the Databricks documentation. # Common expression patterns Source: https://docs.prophecy.ai/data-analysis/gems/visual-expression-builder/use-the-visual-expression-builder Build conditional logic, filters, and reusable expressions with the visual expression builder Use the visual expression builder to create conditional logic, combine multiple filter conditions, build reusable expressions with parameters, and define complex business rules without writing SQL manually. Common use cases include: * categorizing records based on conditions * building multi-condition filters with `AND` and `OR` * creating derived columns * handling null values * using parameters in expressions * creating reusable business logic across pipelines This page explains common expression patterns you can create with the visual expression builder. ## Common expression examples | Goal | Example pattern | | :---------------------------------- | :------------------------------------------ | | Categorize records using conditions | `WHEN revenue < 1000000 THEN 'Low Revenue'` | | Combine multiple conditions | `Amount > 100000 AND Region = 'APAC'` | | Replace null values | `COALESCE(region, 'Unknown')` | | Build nested logic | Multiple `WHEN` clauses with `ELSE` | | Filter rows using grouped logic | Nested `AND` / `OR` conditions | | Use runtime values in expressions | Configuration variables and parameters | ## Expressions that return values vs conditions Some expressions return calculated values, while others return `true` or `false`. For example: * Filter conditions must evaluate to `true` or `false`. * Reformat expressions can return calculated values such as text, numbers, or dates. Understanding the expected output type can help prevent validation and runtime errors. ## Expression-building features The following table describes options for the visual expression builder. | Feature | Description | | ---------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Comparison | Lets you establish relationships between two simple expressions connected by an operator. This mode helps you perform comparisons or existence checks on your data that always evaluate to true or false. | | Grouping | Combine multiple comparison expressions into groups using logical operators like `AND` and `OR`. This structure enables you to express intricate business logic in a visual format. | | Parameters | Enables you to include variables in your expressions that may vary at runtime. Parameters only show up in the **Configuration Variables** of the visual expression builder after you have [created them](/data-analysis/development/parameters/parameters) at the pipeline level. | ## Examples ### Create conditional logic for derived columns Let's say you want to stratify accounts based on their annual revenues. Each condition we set up is limited to one comparison. This example combines conditional logic with comparison operators. Reformat gem using Comparison mode #### Create a new conditional column To set up the comparison expressions to match the image above: 1. In the Reformat gem, under **Target Column**, click **Select Column**. 2. Give the column the name `stratify_by_revenue`. 3. Click **Select expression > Conditional**. A `WHEN` clause appears. #### Configure the WHEN clause 1. For `WHEN`, click **Select expression > Function**. 2. Select **Data type cast**, which converts a value of one data type into another data type. 3. Select **Throw error on failure** to ensure the pipeline doesn't run if the type cast fails. 4. Click **Select expression > Column** and select `ANNUALREVENUE`. 5. Click **Select data type > Float** to convert the column to a Float type. 6. Click **Select operator** and select `less than`. 7. Click **Select expression > Value** and enter `1000000` as the value. #### Configure the THEN clause 1. For `THEN`, click **Select expression** and select **Value**. Enter `Low Revenue` as the value. 2. Click `+` on the next line and select **Add CASE** to add another `WHEN` clause. 3. Repeat steps 3 to 8 to set up the rest of the comparison expressions. 4. Click `+` on the next line and select **Add ELSE** to add an `ELSE` statement. 5. Click **Select expression** and select **Value**. Enter `Unknown` as the value. This conditional expression will categorize your accounts based on revenue thresholds, making it easier to perform segment-specific analysis and reporting. When the pipeline runs, each account will be assigned to the appropriate revenue category based on the conditions you've defined. ### Build multi-condition filters with AND/OR logic When filtering data, you often want the output data to meet multiple criteria. You can use Grouping for this by creating multiple `AND` and `OR` statements. Assume you have a dataset where you want to filter for the following: * Total expected revenue that `is not null` * Total amounts that are greater than `100000` * Latest closed quarters that equals `2023Q2` or `2024Q2` Filter gem using Grouping mode You can have any number of groups and nestings (a group within a group). You can also always change the grouping conditions between `AND` and `OR`. #### Set up base filter conditions To set up the grouping expressions to match the image above: 1. After creating the Filter gem, click **Add condition**. An option to Select expression appears. 2. Click **Select expression > Column**. 3. Select `TOTAL_EXPECTED_REVENUE` from the list. 4. Click **Select operator** and select `is not null`. 5. Click **+ Add Condition** to add another condition expression. 6. Click **Select expression > Column**. 7. Select `TOTAL_AMOUNT` from the list. 8. Click **Select operator** and select `greater than`. 9. Click **Select expression > Value**. 10. Enter `100000` as the value. #### Add grouped `OR` condition 1. Click **Add Group**. A grouped expression row appears. 2. Click **Select expression > Column**. 3. Select `LATEST_CLOSED_QTR` from the list. 4. Click **Select operator** and select `equals`. 5. Click **Select expression > Value**. 6. Enter `2023Q3` as the value. 7. Click **+ Add Condition** and repeat steps 2 to 6 to set up the other `OR` condition. This complex filter will return only high-value opportunities from specific quarters that have valid expected revenue values. By combining AND and OR conditions in this way, you can create precise data subsets that match your exact business requirements. ### Create reusable expressions with parameters When you use a [pipeline parameter](/data-analysis/development/parameters/parameters) in a visual expression, you can manipulate the value of that parameter using different configs at runtime. Let's review an example that leverages an array parameter in a Filter gem. Imagine that you want to filter an `Orders` dataset based on the region where the order was placed. Specifically, you only want to keep rows where the region is included in the array parameter. #### Create an array parameter First, you'll set up a `region` parameter, which will be an array of strings that includes a subset of regions. 1. Open your project and select **Parameters** in the header. 2. Click **+ Add Parameter**. 3. Name the parameter `region`. 4. Select the **Type** and choose **Array > String**. 5. Click **Select expression > Value**. 6. Type `AMER` and click **Done**. 7. Select `+` to add another string to the array. 8. Type `APAC` and click **Done**. 9. Now, click **Save**. Create string array #### Use the parameter in an expression Now, you'll use the parameter in an expression inside a Filter gem. 1. Create and open the Filter gem. 2. Remove the default `true` expression. 3. Click **Select expression > Function** and select `array_contains`. 4. In the **array** dropdown of the function, click **Configuration Variable** and select the `region` parameter. 5. In the **value** dropdown of the function, click **Column** and select the order region column. Filter using array The output of this gem will only include rows where the order region matches at least one value in the `region` array. When you run the pipeline interactively, it will use the values of the default array that you set up in the previous section. ## Validate your expressions Run the pipeline up to and including the gem with your expression, and observe the resulting data sample. To do so, click the **play** button on either the canvas or the gem. Once the code has finished running, you can verify the results to make sure they match your expectations. You can explore the result of your gem in the [Data Explorer](/data-analysis/development/runs/data-explorer/data-explorer). ## Common issues ### Filter expressions failing validation Filter conditions must evaluate to `true` or `false`. For example: * `Amount > 1000` is valid. * `CASE WHEN Amount > 1000 THEN 'High' END` is not valid as a filter condition because it returns text instead of a boolean value. Use conditional expressions to create derived columns in gems like [Reformat](/data-analysis/gems/prepare/reformat), and use boolean comparisons in [Filter](/data-analysis/gems/prepare/filter) gems. ### Conditional logic not returning expected results Verify that: * conditions are evaluated in the expected order. * grouped `AND` and `OR` conditions are structured correctly. * all possible cases are handled with an `ELSE` condition when appropriate. ### Null values causing unexpected behavior Some expressions return `NULL` when input values are `NULL`. Use functions such as `COALESCE()` to provide default values when needed. ### Expression validation errors Validation errors can occur when: * data types do not match. * required function arguments are missing. * column references are incorrect. * expressions return an unexpected value type. ### Nested field references not resolving correctly When working with nested or structured data, verify that: * the correct field path is selected. * the referenced field exists in the input schema. * the expression uses the expected nested structure. ### Multi-condition filters returning unexpected rows When combining `AND` and `OR` conditions: * use grouping to control evaluation order. * verify that conditions are nested correctly. * test expressions incrementally to confirm the output. ## Tips Here are some additional tips to keep in mind when using the visual expression builder: * The expression dropdowns support search. * Each argument of your function is another expression since you have the same expression options to choose from. * You can drag and drop your comparison expressions to rearrange them. * Just as with conditions, you can also drag and drop your grouping expressions to rearrange them. * You can delete individual expressions, conditions, and groupings by clicking the trash icon at the end of the rows. # Variant data type Source: https://docs.prophecy.ai/data-analysis/gems/visual-expression-builder/variant-schema About variant data types and how to update their schema A variant data type is an array of values with more than one data type and provides flexibility to handle diverse and unstructured data from multiple sources without enforcing a rigid schema. This adaptability accommodates data that evolves or comes from different environments, enabling seamless integration and storage. You can use Prophecy to convert your variant data into flat, structured formats to make them easier to understand and use for analytics. This helps you determine the data types of each value in your Snowflake array or object. In Prophecy, you can do the following with variant data: * Infer the variant schema * Configure the parsing limit for inferring the column structure * Use a nested column inside of the visual expression builder ## Infer and edit variant data types Prophecy does not store variant data types within the table definition. Each row can vary in data types, which makes them difficult to infer and use. Fortunately, you don't have to infer the schema yourself. You can use the column selector inside of your gems to automatically infer the variant data type, explore the multi-type variant structure, and later select a nested column to use in your transformations. To automatically infer the variant data type: 1. Open a gem that uses a variant column input, such as the [FlattenSchema gem](/data-analysis/gems/prepare/flatten-schema). 2. Click the **Variant** dropdown, and click **Infer Schema**. Prophecy automatically detects and identifies the variant data types in your input data. Schema and column selector Prophecy caches the inferred schema so you can use it again in the future whenever you reopen the model, gem, or another gem connected to the same input port. To see the last time your variant data type was inferred, see the box at the bottom of the column selector. To refresh the schema, click **Infer Schema** again. ### Edit the variant data type schema If Prophecy missed certain schema cases while sampling the records, you can make edits yourself. To edit the variant data type schema: 1. Click the **Variant** dropdown, and click **Edit Schema**. 2. Use the data type dropdowns to manually choose the data type of each nested schema. Edit schema view ## Variant sample setting When Prophecy infers the variant data type, it samples the records to identify all potential iterations of keys and values within the schema. The default number of records that Prophecy parses to understand the nested data schema is 100. To update this limit: 1. Click on `...` at the top of the page. 2. Click **Development Settings**. 3. Put in the number of records you want Prophecy to parse to understand the nested data schema. We recommend that you increase the limit for small structures, or decrease it for larger ones. Variant sampling setting This setting does not rely on the ratio of the data since that would require a complete count of the data records. ## Use a nested column in an expression In the column selector, you can add a nested column by clicking **Add Column** when you hover over the input field name. Add column When you add a column nested within a variant data type, Prophecy automatically generates several fields according to the following rules: | Field | Rule | | ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Column name | The column name matches the input field name, and is prefixed with the parent field path. If there's a conflict, Prophecy appends numbers starting with `_0` until it becomes unique.
For example, if the column name `customers_name` already exists, Prophecy might name the new field `customers_name_0`. | | Expression | The expression represents the full path to the selected field, and uses existing flattened subpaths. | | Data type | The data type is automatically `CAST` to the closest inferred type. | ### Default casting Prophecy automatically adds a `CAST` to any column you add from a nested type. By default, the column is cast using the standard `CAST(x AS y)` syntax. In some cases, a path within a variant data type may hold different value types across rows. For example, a dataset can contain different data types, such as integer, object, and boolean for each row's value key. Prophecy supports this scenario by presenting each detected data type for a given key, array, or object as a separate item in the column selector. When you add one of those columns to the expression, Prophecy uses explicit casting, which may error out if the cast is not possible. You can change this behavior by using `TRY_CAST`, which returns `null` if the cast is not possible. # Visual expression builder Source: https://docs.prophecy.ai/data-analysis/gems/visual-expression-builder/visual-expression-builder About the visual expression builder Use the visual expression builder to write complex SQL expressions without worrying about syntax. The visual expression builder can help you to better understand the relationships between different functions and their arguments. You can use visual expressions in gems and data tests. To understand how to build expressions with the visual expression build, see the [reference guide](/data-analysis/gems/visual-expression-builder/visual-expression-builder-reference). ## Build with Copilot You can use Copilot from the visual expression builder for additional help. Whether you're exploring functions, learning what's possible, or writing expressions with prompts, Copilot supports your workflow. Copilot in the visual expression builder ## Code view To view the SQL expressions generated by the visual expression builder, you can switch to the Code view of a gem or of the project. If you update any expressions in the Code view, they will be converted back to visual expressions in the Visual view. Code Expression Builder You can also ask Copilot to generate SQL expressions directly in the Code view. The SQL dialect for expressions depends on the SQL warehouse connection of the fabric. This is because the project code must be compatible with the execution environment it runs on. Because of this, expressions may look different across projects that run on Databricks versus Snowflake, for example. ## What's next To continue developing with the visual expression builder, see the following pages: # Visual expressions reference Source: https://docs.prophecy.ai/data-analysis/gems/visual-expression-builder/visual-expression-builder-reference visual expression builder reference This page contains a reference of the different visual expression builder components, which include the expression options, operator options, and data types. ## Expression options The visual expression builder supports the following expression options: | Option | Description | | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Column | Select an input column from your source tables.
You can view all of the available input columns from the dropdown menu. | | Value | Enter any value.
  • If you enter a string value, it'll be considered as a string within quotes.
  • If you enter a number, it'll be considered a numerical value, but you can click Check to read value as string.
  • The same applies to boolean values. For example, if you enter True, it'll be considered a boolean unless you select Check to read value as string.
| | Function | Choose a function from a list of all supported function category groups.
The list displays each function's description, including mandatory arguments. You can optionally add additional arguments if applicable. | | Data type cast | Cast a variant column into its appropriate data type.
  • Use TRY\_CAST to avoid errors. On failure, it sets the value to null.
  • Note: For Snowflake, TRY\_CAST is only supported on string data types.
| | Conditional | Use a conditional `WHEN` clause.
  • Within WHEN, use a comparison expression.
  • Within THEN, use a simple expression.
  • You can add multiple CASES but only one ELSE.
  • ELSE also uses a simple expression.
  • You can also add IF, ELSEIF, or FOR conditions.
  • FOR uses a variable name and expression value.
  • IF and ELSEIF are comparisons.
| | Configuration Variable | Use a variable from the list of your [pipeline parameters](/data-analysis/development/parameters/parameters) or project variables. | | Incremental | Use advanced dbt configurations. | | Custom Code | Write your own custom code for expressions not supported by the visual expression builder.
As you type, suggestions will be provided. | ## Data types | Category | Supported Types | | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Basic |
  • `Boolean`
  • `String` - String / Varchar
  • `Date & time` - Date / Datetime / Timestamp / Timestamp NTZ
  • `Number` - Integer / Long / Short
  • `Decimal number` - Decimal / Double / Float
| | Other |
  • `Binary`
  • `Byte`
  • `Char`
  • `Calendar interval` / `Day time interval` / `Year month interval`
  • `Null`
  • `Variant`
| ## Comparison operators | Operator | Description | | ----------------------- | ------------------------------------------------------ | | `equals` | Checks if two values are equal. | | `not equals` | Checks if two values are not equal. | | `less than` | Checks if a value is less than another. | | `less than or equal` | Checks if a value is less than or equal to another. | | `greater than` | Checks if a value is greater than another. | | `greater than or equal` | Checks if a value is greater than or equal to another. | | `between` | Checks if a value lies between two others. | ## Existence checks | Operator | Description | | ------------- | ------------------------------------------- | | `is null` | Checks if a value is null. | | `is not null` | Checks if a value is not null. | | `in` | Checks if a value exists in a given list. | | `not in` | Checks if a value does not exist in a list. | ## Boolean predicates | Type | Predicates | | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------ | | Unary |
  • `Exists` (in subquery)
  • `In`
  • `Is null`
| | Binary |
  • `Between`
  • `Equality`
  • `Less than`
  • `Less than or equal`
  • `Greater than`
  • `Greater than or equal`
| | Groups |
  • `Not`
  • `And`
  • `Or`
| # Concepts for Data Analysts Source: https://docs.prophecy.ai/data-analysis/getting-started/concepts Core concepts for data analysts working with Prophecy ## AI Prophecy provides a suite of AI capabilities to help you build data pipelines faster and more efficiently. Use AI to build pipelines, fix issues, generate documentation, and more. * **Agent**: The [Prophecy Agent](/data-analysis/ai/agent/agent) specializes in building pipelines, adding transformations, and creating complete data workflows. * **Copilot**: The Prophecy Copilot specializes in making quick fixes, suggesting expressions, and other one-click tasks. Prophecy is optimized for Google Chrome. ## Projects [Projects](/data-analysis/development/projects/create-project) serve as your primary workspace in Prophecy for building data pipelines. They are containers that organize related data transformations, tests, and schedules in one place. ## Pipelines [Pipelines](/data-analysis/development/pipelines/data-analysis-pipelines) are essential components in data processing workflows, enabling the automated movement and transformation of data. They define a sequence of steps that extract data from a source, process or transform it, and load it into a destination system. ## Gems [Gems](/data-analysis/gems/gems) are visual, modular components that represent data transformation logic in Prophecy. Each gem encapsulates a specific operationβ€”such as reading data, filtering rows, or aggregating valuesβ€”and automatically generates the corresponding code in your project's language. Gems are designed to connect together in a pipeline where data flows from one gem to the next. ## Studio All of your pipeline development happens in the [Studio](/data-analysis/development/studio/studio). The Studio is a visual interface that allows you to build your pipelines. You can add and connect gems to the canvas to build your pipeline. When you open a project in Studio, Prophecy opens the last saved state when one is available. If there is no saved state, Prophecy opens a project landing page where you can open existing entities or create new ones. ## Analyses [Analyses](/data-analysis/analysis/overview) are dashboards that let you visualize pipeline outputs and surface key insights on pipeline data. You can make these interactive using [pipeline parameters](/data-analysis/development/parameters/parameters). ## Fabrics [Fabrics](/data-analysis/environment/fabrics/prophecy-fabrics) are Prophecy-specific entities that define the execution environment for your pipelines, including SQL warehouse connections, data ingress/egress connections, and secrets. Fabrics are automatically provisioned for users on the Professional Edition. ## Teams [Teams](/data-analysis/administration/management/teams/teams) represent groups of users who collaborate on projects and share access to resources. When you create a project or a fabric, you assign it to a team. All users in that team will have access to the relevant project or fabric. # Use Professional Edition Source: https://docs.prophecy.ai/data-analysis/getting-started/professional-edition A managed, AI-first analytics environment for data analyst teams Prophecy's Free and Professional Editions are designed for data analyst teams that want to work with AI-driven analysis without managing data infrastructure. Prophecy provides and manages the underlying resources, metered by credits. The Free Edition is functionally identical to the Professional Edition, but is limited to 5 credits per month and a single user per plan. Sign up for the Free Edition to get started. You can upgrade to Professional Edition after you sign in for the first time. Coming from Alteryx or prefer to build pipelines directly? Prophecy for Business gives you a full visual pipeline editor and SQL editor alongside the Agent. See [Prophecy for Business](/data-analysis/getting-started/concepts) to compare. ## What you can do Professional Edition centers on a conversational workflow: attach data, describe what you want to understand, and the Agent generates pipelines and analyses in response. From the home page you can: * Attach a data file and start an analysis with a prompt or a common task. * Resume recent analyses and chat threads from your project. * Add datasets, pipelines, analyses, and skills to your project from the project panel. The Agent handles pipeline generation in the background. You interact primarily through chat and the analysis surface rather than the pipeline canvas directly, though pipelines remain visible and inspectable within your project. Professional Edition Home Page ## Get started 1. Sign up for the [Free Edition](https://app.prophecy.ai/professional/). 2. Sign in and upgrade to Professional Edition when ready. 3. From the home screen, select **Attach Data** to upload a file. 4. Enter a prompt or choose a common task β€” **Summarize my dataset**, **Find patterns or trends**, or **Clean up my data**. 5. Select **Get Started**. 6. Choose **Process as Table** to run transformations against your data, or **Attach as Reference** to make the file available for context. After you select **Get Started**, the view splits into two panels. The left panel shows the Agent's response. The Agent suggests analysis questions based on what it finds. You can ask a follow-up question directly, or select one of the Agent's suggestions to go deeper. The conversation continues here as you refine your analysis. The right panel shows your project. Your attached file appears under **Browse Project**. Recent chat threads are listed under **Recent Activity**. From here you can also select **+ Add New** to add a dataset, pipeline, analysis, or skill to the project as your work grows. Professional project page For some requests the Agent responds inline without generating a pipeline. For others it builds and runs a pipeline in the background. When a pipeline is generated you can inspect it at any time from the project panel. ## What's different from Prophecy for Business Professional Edition runs on Prophecy-managed DuckDB. You do not configure fabrics, connections, or secrets β€” there is no infrastructure to set up. The tradeoff is scope. Professional Edition does not currently support connecting to external data platforms such as Databricks, Snowflake, or BigQuery. Other differences from Prophecy for Business: * No versioning. * No fabric, connection, or secrets management. * No scheduling or publishing. ## When to move to Prophecy for Business Consider moving to Prophecy for Business when you need to: * Connect to Databricks, Snowflake, or BigQuery. * Schedule and deploy pipelines to production. * Version and branch project work. * Manage connections and secrets at the team or organization level. Your projects and analyses carry over. Professional Edition is designed as a starting point, not a ceiling. # Quickstart for Data Analysts Source: https://docs.prophecy.ai/data-analysis/getting-started/quick-start Use the Agent to develop a simple pipeline *Estimated time: 30 minutes* Build a data pipeline using Prophecy Agent. This tutorial walks you through exploring data, creating visualizations, and building transformations using natural language prompts. Follow the steps below to build a patient analytics pipeline. Importing from Alteryx or other platforms? Start with [Import Workflows into Prophecy](/import-tool/). ## Prerequisites To complete this quickstart, you need a [Prophecy fabric](/data-analysis/environment/fabrics/prophecy-fabrics) that uses Prophecy In Memory or Databricks as the compute engine. Prophecy automatically creates a compatible fabric for Free and Professional Edition users. If you do not have any existing fabrics, [you'll need to create one](/data-analysis/environment/fabrics/prophecy-fabrics). ## Set up a new project First, you need to create the project where you will build your pipeline. You'll also need to add data to the project for this quickstart. 1. Click on the **Create Entity** button in the left navigation bar. 2. Hover over the **Project** tile and click **Create**. 3. Give your project a **Name**, such as `Prophecy_Quickstart`. 4. Under **Team**, select your personal team. (It will match your user email.) 5. Under **Select Template**, choose **Prophecy for Analysts**. 6. Click **Complete**. Prophecy will open your new project in [Studio](/data-analysis/development/studio/studio). For a new project with v4 AI enabled, Agent chat opens maximized by default. 1. Select the default fabric or your own fabric. 2. Click **Save**. This attaches the project to the fabric. You can switch fabrics at any time by clicking the fabric selector in the project header. 1. Minimize Agent chat to show the project landing page. 2. On the project landing page, click **Create Pipeline**. 3. For the **Pipeline Name**, enter `patient_analytics`. 4. Leave the default **Directory Path** of `pipelines`. Prophecy saves your compiled pipeline code in this folder of the project repository. 5. Click **Create**. This opens the pipeline canvas for your new pipeline. ## Add data to your project Next, you'll add some data to the project so the Agent can find and transform it. Load some data into the project as a Seed: 1. Open the **Source/Target** gem category. 2. Click **Table**. This adds a new [Table gem](/data-analysis/gems/source-target/source-target) to the canvas. 3. Hover over the gem and click **Open**. 4. Select **+ New Table**. 5. For the **Type and Format**, choose **Seed**. 6. Name the seed `patients_raw_data`. 7. For the **Seed path**, choose **seeds**. Prophecy saves your seed file in this folder of the project repository. 8. Click **Next**. 9. In the **Properties** tab, paste the following data. ```csv patients_raw_data.csv theme={null} patient_id,first_name,last_name,city,state,county,age,admission_date,diagnosis,treatment_cost 1001,John,Smith,Boston,MA,Suffolk,45,2024-01-15,Hypertension,1250.00 1002,Sarah,Johnson,Boston,MA,Suffolk,32,2024-01-18,Diabetes,2100.50 1003,Michael,Williams,Cambridge,MA,Middlesex,58,2024-01-20,Heart Disease,3500.75 1004,Emily,Brown,Boston,MA,Suffolk,29,2024-01-22,Asthma,850.25 1005,David,Jones,Worcester,MA,Worcester,67,2024-01-25,Hypertension,1450.00 1006,Jessica,Garcia,Boston,MA,Suffolk,41,2024-02-01,Diabetes,2200.00 1007,Christopher,Miller,Springfield,MA,Hampden,53,2024-02-05,Heart Disease,3800.50 1008,Amanda,Davis,Boston,MA,Suffolk,35,2024-02-08,Asthma,920.75 1009,James,Rodriguez,Cambridge,MA,Middlesex,62,2024-02-10,Hypertension,1320.00 1010,Lisa,Martinez,Worcester,MA,Worcester,48,2024-02-12,Diabetes,2150.25 ``` 10. Click **Next**. 11. Click **Load Data** to preview the data in tabular format. 12. Click **Save**. To materialize the seed data into the SQL warehouse, run the pipeline once. To do so, click the play button in the bottom right corner of the canvas. Prophecy should automatically index the table when you save the seed. This allows the Agent to discover and use the seed data. If you have trouble finding the table in later steps, you can also manually reindex the [knowledge graph](/data-analysis/ai/knowledge-graph/knowledge-graph). 1. Open the **Environment** tab in the left sidebar. 2. Below your connections, you'll see a **Missing Tables?** callout. 3. Click **Refresh** to trigger the knowledge graph indexer. You'll see a progress bar in the callout indicating the indexer is running. Once it's complete, the callout will disappear. Verify the table was indexed by checking the **Environment** tab in the left sidebar. ## Explore the data Now that you have data in your project, you can explore it using the Agent. If Agent chat is minimized, maximize it to continue working in chat. This is where you'll interact with the Agent. Ask the Agent to search for the seed data table. Enter the following prompt in the chat: ``` Find datasets with information pertaining to hospital patients ``` The Agent returns: * A short list of relevant datasets with descriptions. * A full list of matching datasets. * The option to add datasets directly to your pipeline on hover. 1. Click on `patients_raw_data` (the data you uploaded) in the chat to open a detailed preview dialog where you can: * View the table location. * Examine the schema and column structure. * Preview sample data. * Review data profiles. * Open an **Explore** session for dataset-specific queries. 2. Close the preview dialog to return to the chat. If the Agent doesn't find your table, verify that the table was indexed by checking the **Environment** tab in the left sidebar. Additionally, if you have other patient data in your warehouse, the Agent may return other datasets that match your query instead. Request specific data samples to validate your understanding. Enter the following prompt, replacing `@patients_raw_data` with your actual table path: ``` Provide sample data from @patients_raw_data showing only patients from Boston ``` The Agent returns: * A table with the requested data. * An option to preview the table for detailed examination. * An option to add the table to your pipeline. * SQL execution logs showing the query used. Verify the results show 4 patients: John Smith, Sarah Johnson, Emily Brown, and Jessica Garcia, all from Boston. Generate charts and insights from your data. Enter the following prompt: ``` Visualize number of patients per city in @patients_raw_data ``` The Agent returns: * An embedded chart in the chat showing patient counts by city. * An option to preview the table for detailed examination. * An option to add the table to your pipeline. * SQL execution logs showing the query used. Click **Preview** to access: * **Visualization tab**: View larger charts and download charts as images. * **Data tab**: Examine underlying data and download data as JSON/Excel/CSV. The chart should show Boston with 5 patients, Cambridge with 2, Worcester with 2, and Springfield with 1. ## Build your pipeline Now, you'll build your pipeline transformations using the Agent. Keep the **Visual** view open (rather than the **Code** view) to see updates in real-time as the Agent adds or modifies gems on the canvas. Describe the transformation you want to perform. Enter the following prompt: ``` Transform @patient_records to show total number of patients per county ``` Replace `@patient_records` with the actual gem label from your canvas if it differs. The Agent will: * Add the appropriate gem(s) to your pipeline canvas. This prompt should produce an [Aggregate](/data-analysis/gems/transform/aggregate) gem. * Execute the pipeline, generating data samples that you can review. * Provide a description of the changes made. * Show options to inspect, preview, or restore changes. * Display SQL execution logs. The output should show 3 counties: Suffolk with 5 patients, Middlesex with 2, Worcester with 2, and Hampden with 1. Each change that the Agent makes can be viewed in the [project version history](/data-analysis/development/versioning/version-control#show-version-history). You can revert the changes at any time. To understand what the Agent built: 1. Click **Inspect** on the transformation response. 2. Review the configuration panel starting with the first modified gem (highlighted in yellow). 3. Use the **Previous** and **Next** buttons to navigate through modified gems. 4. Examine input and output data to verify the transformation produces expected results. This helps you: * Understand the Agent's approach. * Verify that the transformation logic matches your expectations. Build a more complex transformation. Enter the following prompt: ``` Calculate average treatment cost per county from the previous step ``` The Agent adds another transformation that calculates the average cost. This demonstrates how you can chain transformations together. Verify the results show average costs for each county. Suffolk should have an average around `1,544`, Middlesex around `2,410`, Worcester around `1,800`, and Hampden `3,800.50`. After building your pipeline, ask the Agent to save the final output. Enter the following prompt: ``` Save the final output of this pipeline as a table ``` The Agent adds a [Table gem](/data-analysis/gems/source-target/source-target) to the end of your pipeline. When you run the pipeline, the Table gem writes the data to your default database and schema defined in your fabric, allowing you to persist results. ## Explore further Try these additional tasks to extend your pipeline. Enter the following prompt in the chat: ``` Filter the source data to only include patients over 50 years old ``` Create a second Seed file named `county_info` with the following content: ```csv county_info.csv theme={null} county,population,region Suffolk,800000,Eastern Middlesex,1600000,Eastern Worcester,830000,Central Hampden,470000,Western ``` Then, prompt the Agent: ``` Join @l0_raw_patients with @county_info on county ``` The Agent will infer join keys, but you can also specify them explicitly. Enter the following prompt in the chat: ``` Add a column that calculates days since admission using the current date ``` ## Connect your own data This quickstart uses a Seed file as the source data. When you start building your own pipelines, you'll likely want to use your own data from files or external systems. You can do this by: * Uploading files directly from your local filesystem using the [upload file](/data-analysis/gems/source-target/table/upload-files) feature. * Ingesting data from external systems using [connections](/data-analysis/environment/connections/connections). ## Sample prompts reference Use these prompts as templates for your own pipelines. | Task | Prompt Example | | -------------- | ----------------------------------------------------------------------------------------------------- | | Find data | `Find datasets containing customer information` | | Sample data | `Show me 5 random records from @sales_data` | | Filter data | `Filter to only include orders from 2024` | | Transform data | `Calculate total revenue as quantity * price` | | Join data | `Join the orders and customers tables` | | Parse data | `Extract the fields from json_data as columns` β€” This works best if you provide a sample JSON object. | | Clean data | `Remove rows where email is null` | | Aggregate data | `Group by region and calculate average sales` | | Visualize data | `Create a bar chart of monthly sales` | | Save results | `Save the final output as a table` | ## Tips **How to do it**: Instead of `Clean the data` β†’ Try `Remove duplicate customer records and fill null values in the email column` **Why it helps**: Reduces ambiguity so the Agent applies the correct operations. **How to do it**: Break complex transformations into smaller requests. For example: `Aggregate orders by customer` β†’ `Join with customers` β†’ `Filter for orders made in 2024` **Why it helps**: Improves reliability and makes it easier to debug or adjust each step. **How to do it**: After a response, click **Inspect** and use **Previous/Next** to review highlighted gem configuration and output **Why it helps**: Ensures the transformation matches expectations before you continue. **How to do it**: Use chat to scaffold transformations, then switch to the **Visual** canvas to fine-tune or add gems **Why it helps**: Combines speed (AI) with precision and control (visual editor). ## Troubleshooting **Probable cause**: [Knowledge graph](/data-analysis/ai/knowledge-graph/knowledge-graph) is out of date **How to fix**: [Reindex your fabric connection](/data-analysis/ai/knowledge-graph/indexer) so the knowledge graph includes the table **Probable cause**: Chat context isn't relevant anymore **How to fix**: Start a new chat. **Probable cause**: Seed data wasn't created correctly **How to fix**: Verify the table exists and contains the expected 10 rows of data # Ad Hoc & Document-Based Workflows Source: https://docs.prophecy.ai/data-analysis/getting-started/structured-finance/adhoc-document-workflows Extract tables from a prospectus and reconcile them against your harmonized tape You can build fully custom workflows against source documents, such as a prospectus, a servicer report, or other reference document, by uploading them and asking the Agent to use their contents. A prospectus comparison is one common example. ## Example: extract from a prospectus Upload a prospectus (for example, a 424B5 filing) and ask the Agent to extract a specific table β€” referencing it by page number β€” and compare it against your existing Strats output. Ask the Agent to share its plan before it executes. ``` "On page 66, we have a distribution of receivables as of cutoff date by remaining principal balance. Can you extract this table and compare it with the one we have in our strats?" ``` upload prospectus The Agent extracts the table and can often self-correct even if the page number given is slightly off. ### Align bucket structures Prospectus tables and your Strats output may bucket values differently. To compare them directly, align your Strat buckets to match the prospectus structure β€” the best practice is to modify the template directly rather than reconciling mismatched buckets after the fact. First-time runs of a custom ad hoc workflow typically need some tuning. Treat the first pass as a baseline to refine, not a final result. ### Generate a shareable report Once the comparison is built, ask the Agent for a shareable Markdown report combining the work β€” for example, `create a proper MD with a companion report I can share.` ### Example output **TAOT 2026-A: Prospectus vs. Tape Reconciliation Report** * Deal: Toyota Auto Receivables 2026-A Owner Trust * Prospectus Filing: 424B5 (SEC EDGAR) * Tape Source: Harmonized CDM (`output_preview`) | Metric | Prospectus | Tape | Variance | | ----------------- | --------------- | --------------- | ---------------- | | Loan Count | 65,867 | 65,867 | 0 | | Aggregate Balance | \$1,995,353,676 | \$2,034,742,702 | +\$39.4M (+2.0%) | | Average Balance | \$30,294 | \$30,892 | +\$598 (+2.0%) | The report's executive summary calls out the variance directly, noting, for example, that the tape shows higher balances concentrated in the upper balance tiers relative to the prospectus disclosure. # Export & Sharing Source: https://docs.prophecy.ai/data-analysis/getting-started/structured-finance/export-sharing Export analyses as PDF or Excel, share project snapshots, and update results through the agent ## Exporting Click the three dots in the upper right-hand corner of an analysis to export it as **PDF** or **Excel**. The export is saved in the system as a read-only file backed by the underlying pipeline. * Excel and PDF exports appear under **Browse Project > Files**. * You can open either directly in Prophecy. Neither is editable. ## Updating an export To change anything in an exported analysis, return to the agent and make the change there β€” don't edit the export directly. The Excel file regenerates automatically. For example, ask the agent to `add a distribution to this Strats based on vehicle_model`, then follow up with a refinement like "top 20." The agent rebuilds the table and pipeline and reruns the analysis. ## Sharing The **Share** button at the top of the page shares a project snapshot with other users. ## Table-level actions Any table in an analysis can be: * Sorted * Reordered by dragging columns * Downloaded as JSON, table, or Excel via the **Download** button # Overview Source: https://docs.prophecy.ai/data-analysis/getting-started/structured-finance/overview What Structured Finance does and how the workflows fit together Structured Finance is an agent-driven workflow for onboarding, harmonizing, and analyzing ABS collateral tapes. It takes a raw loan or lease tape through to a reviewed, standardized dataset, then supports stratification, cross-deal comparison, and prospectus-based analysis on top of it. ## Supported deal types Structured Finance works with Auto ABS deals, including GMCAR, SDART, BMWOT, TAOT, CALT, and EART. ## Workflow Upload a raw tape and harmonize it against a Common Data Model (CDM). The agent maps source fields to the target schema, runs data quality checks, and surfaces mappings for review. Run stratification analyses on the harmonized tape using built-in or custom templates. Every result traces back to an inspectable pipeline. Export results as PDF or Excel, or share a project snapshot. Changes made through the agent regenerate exports automatically. Use built-in AI skills β€” ad-hoc stratification queries, cross-deal comparison, CDM authoring, and spec generation β€” triggered by natural-language prompts. Each phase is covered in its own page: * [Tape Cracking: Onboarding & Mapping](/data-analysis/getting-started/structured-finance/tape-cracking) * [Stratifications & Pool Selection](/data-analysis/getting-started/structured-finance/stratifications-pool-selection) * [Export & Sharing](/data-analysis/getting-started/structured-finance/export-sharing) * [Use Skills](/data-analysis/getting-started/structured-finance/skills) * [Ad Hoc & Document-Based Workflows](/data-analysis/getting-started/structured-finance/adhoc-document-workflows) # Use Skills Source: https://docs.prophecy.ai/data-analysis/getting-started/structured-finance/skills Use built-in AI skills for stratification, cross-deal comparison, CDM authoring, and spec generation Structured Finance includes built-in *skills*: reusable capabilities that combine a set of prompts and domain logic into a single task the agent can carry out on request. For example, the `adhoc-strats` skill turns a plain-language question like "show me top 10 vehicle models by balance" into a full stratification query, without you having to build the aggregation yourself. You can ask the Agent what's available: ``` what Structured Finance skills are available? ``` and it will return a list of built-in skills. You can also [add your own skills](#add-a-new-skill) or have the Agent create one for you. ## Built-in skills | Skill | Trigger | Purpose | | --------------- | ------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------- | | `harmonization` | "harmonize", "crack a tape", "map to CDM" | Map raw source data to a target CDM schema, edit mappings, fix DQ failures | | `adhoc-strats` | "show me", "what's the distribution", "top 10", "weighted average" | Ad-hoc stratification queries on collateral pools β€” FICO, APR, LTV, state, vehicle, delinquency breakdowns | | `asset-comp` | "asset comp", "compare deals", "extract comp", "research note" | Build and query cross-deal comparison tables, extract metrics from prospectuses, generate research notes | | `cdm` | "create CDM table", "add column to CDM", "DQ checks" | Author and manage Common Data Model schemas with data quality rules | | `livespec` | "generate spec", "create documentation" | Auto-generate pipeline and project specification documents | These skills work across all supported Auto ABS deal types (GMCAR, SDART, BMWOT, TAOT, CALT, EART, and so on) and integrate with the harmonization and strats pipelines you're already using. Skills are customizable to your organization's needs. When you confirm use of a skill, the Agent identifies the necessary steps and gems, executes them, and returns traceable results β€” for example, surfacing state or vintage-quarter concentrations to support a pool-cutting decision. ## Get the full cheat sheet for a skill You can ask the Agent for all cheat sheet prompts: ``` adhoc Strats cheat sheet ``` The Agent then returns a list of prompts: result of adhoc Strats cheat sheet then run specific ones by reference: ``` 2.1 "Show me top 5 states by balance with avg FICO and avg LTV" 2.2 "Show me top 10 vehicle models by balance with avg LTV" 2.3 "Break down the pool by origination quarter with avg FICO and APR" 2.4 "Show me any state over 15% of balance, any model over 10%, any vintage over 40%" ``` You can then run a prompt directly by number β€” for example, `run 2.2`. Running `2.2 Show me top 10 vehicle models by balance with avg LTV`, for example yields a result similar to the following: result of running 2.2  Show me top 10 vehicle models by balance with avg LTV ## Ad-hoc stratification (`adhoc-strats`) The `adhoc-strats` skill accepts natural-language prompts for exploring a loan pool. Prompts are organized around common analysis patterns: * **Pool Health** β€” summary stats (contracts, balance, WA FICO/LTV/APR), credit tier splits, new/used breakdowns * **Concentration Risk** β€” top states, vehicle models, and vintages by balance; concentration threshold flags * **Tail Risk** β€” filtering to subprime FICO, high LTV, or combined tail segments * **Stacked Risk** β€” multi-factor filters (e.g., FICO under a threshold AND LTV over a threshold) * **Pricing Adequacy** β€” APR by credit tier, mispriced-risk detection, new vs. used APR spread * **Delinquency & Performance** β€” DPD bucket breakdowns, delinquency by state or credit tier Prompts support threshold tuning (e.g., "under 660" vs. "under 620"), percentage-of-pool framing, and dimension pivots (e.g., "by state" or "by vintage"). Thresholds for what counts as subprime, high LTV, or high APR vary by deal type β€” prime loan vs. subprime loan, loan vs. lease. ## Cross-deal comparison (`asset-comp`) The `asset-comp` skill runs on any of these trigger phrases: * `asset comp` / `add to comp` / `compare to other deals` * `compare this tape` / `extract comp` / `insights on this deal` * `write research note` / `collateral analysis` / `cross-deal comparison` * `GMCAR comp` / `download prospectus` / `refresh comp` ### Operations | Operation | What it does | Example prompt | | ------------ | --------------------------------------------------- | ---------------------------------- | | **extract** | Pull metrics from the current tape into comp format | "extract comp from this tape" | | **add** | Add a deal to the comp table | "add SDART 2025-1 to comp" | | **compare** | Cross-deal comparison against the comp table | "compare this tape to other deals" | | **insight** | Generate insights on the current deal vs. peers | "insights on this deal" | | **research** | Write a research note or collateral analysis | "write research note" | ### Input modes 1. **Harmonized CDM tape** β€” uses your current pipeline's `output_preview` data 2. **SEC prospectus** β€” downloads and parses prospectus PDFs (e.g., "download prospectus for GMCAR 2025-2") ### Example: extracting a comp "extract comp from this tape" returns a preview β€” no changes are made to the comp table until you confirm: #### SDART 2025-4 β€” Comp Extraction (Preview) | Metric | Value | | ------------ | --------------- | | Pool Balance | \$2,034,742,702 | | Loan Count | 65,867 | | WA APR | 5.67% | | WA FICO | 771 | | WA LTV | 104.96% | | New % | 87.98% | Next steps from a preview: * "add to comp" β†’ insert into `cross_deal_comp.csv` * "compare to others" β†’ side-by-side with peer deals * "insights on this deal" β†’ full cross-deal analysis Cross-deal comparison works because every tape is harmonized to the same CDM, which means that metrics line up without additional mapping. The Agent also self-retries on errors encountered during execution. ## Add a new skill Built-in skills cover common workflows β€” harmonization, ad-hoc strats, cross-deal comparison β€” but you can add a skill scoped to your own team's practices. For example, a desk might codify its own tail-risk definitions as a skill, since what counts as "subprime" or "high LTV" varies by deal type. A skill such as `desk-risk-thresholds` could encode your team's specific cutoffs and stress-test segments as a reusable prompt set, so anyone on the desk gets consistent results without re-specifying thresholds each time. To add a new skill: 1. Click the **+**, then select **New Tab**, then select **Skill**. 2. In the dialog, enter a **Skill Name** and click **Create**. 3. On the new page, enter a description and content for the skill. You can also ask the Agent to create a skill for you. For more information on skills, see [Add and use skills](/data-analysis/ai/using-skills) # Stratifications & Pool Selection Source: https://docs.prophecy.ai/data-analysis/getting-started/structured-finance/stratifications-pool-selection Run strats analysis on a cracked tape and inspect how results are derived Once a tape is harmonized, you can run stratification analysis against it using a built-in or custom template. ## Run a Strats template Click **Strats and Pool Selection**, review the template, and click **Use Template**. Custom templates are supported if your deal requires buckets or breakdowns that don't match the built-in set. The Agent runs, applying the **Strats and Pool Selection** to your cracked tape. After applying the template, you get: 1. An analysis (Prophecy's term for dashboard), which opens automatically. 2. A pipeline showing the full transformation logic behind it, all of which are fully inspectable: aggregations, transforms, joins, unions, and so on. To view the pipeline=, click the pipeline icon in the Agent's response. strats template applied ## Inspect an analysis Click **Inspect** on any part of the analysis to send a prompt to the Agent asking it to show how the values were derived. This makes every result fully explainable rather than a black box. Inspect returns: * **Thought summary** β€” a plain-language explanation of what the analysis does and why. For example, a "Current Principal Balance Distribution" workflow groups loan balances by size range to show concentration of value across segments. * **Operations table** β€” the sequence of gems (workflow steps) that make up the analysis: * **Operations** β€” the gem name and type, for example `ref_notional_bucket...` (a Function gem) or `ref_notional_agg` (an Aggregate gem) * **Description** β€” a short summary of what that gem does, such as assigning each loan to a bucket, totaling balances, or formatting output rows Click any gem in the operations table to open that step directly in the workflow, so you can trace exactly how a result was produced. Continue to [Export & Sharing](/data-analysis/getting-started/structured-finance/export-sharing) to export or share the results. # Tape cracking Source: https://docs.prophecy.ai/data-analysis/getting-started/structured-finance/tape-cracking Attach a tape, select a CDM, and let the Agent map source fields to the target schema Tape cracking is the process of harmonizing a raw loan or lease tape against a Common Data Model (CDM). ## Start harmonization Attach the loan tape (CSV or XML) and tell the Agent to begin. For a new tape, you can enter `let's crack this tape`. Click **Get Started**. Prophecy creates a new project and identifies the schema of the source file. Prophecy organizes work into [projects](/data-analysis/development/projects/create-project). You can return to this project later by clicking the **Projects** icon in the side bar. ## Review the source data Before mapping, you should review row profiles and distributions to understand the tape's data. Click **Data Profile** to see detailed column-level information, including value distributions (such as the percentage breakdown across vehicle types). See [Data Profiling](/data-analysis/development/runs/data-explorer/data-profile) for a full reference on distributions and profile metrics. Click **Continue** once you've reviewed the profile. ## Select a CDM Choose a Common Data Model (CDM), such as `auto_ABS_CDM`. You can click a CDM to view its full definition. The CDM you select is part of a domain pack. A *domain pack* bundles the CDM together with the domain-specific knowledge needed to interpret it correctly. auto\_ABS\_CDM is bundled with US auto loan structures, which is why fields like `obligorEmploymentVerificationCode` or `vehicleModelYear` map correctly out of the box. Click **Map Data** to continue. At this point the file is uploaded to a [fabric](/data-analysis/environment/fabrics/prophecy-fabrics). A fabric is a Prophecy entity that contains the connection information needed to connect to external compute and data storage. You don't need to configure one here; it's used automatically as part of this step. ## Agent-driven mapping Prophecy begins harmonization, starting with deterministic mapping. The agent narrates its reasoning as it works through the tape: ``` Mapping employment & obligor fields... - Mapping obligorEmploymentVerificationCode β†’ code 1 = "Not stated/not verified", code 2 = "Stated/not verified"... continuing to next code - Mapping employment_status β†’ code 3 = "Stated/verified", defaulting NULL values to this mapping - Checking seller_subvented_flag β†’ contains '1', '2', or '98'? Evaluating condition - Calculating numberOfObligors β†’ setting to 2 when coObligorIndicator is true, else 1 - Casting seller_loan_id and issuer_loan_number β†’ target type VARCHAR - Computing vehicle_age_months β†’ diffing vehicleModelYear against originationDate - ⚠️ Flagging for review: vehicleModelYear looks like a raw year value (e.g. 2020) β€” verifying calculation formula before finalizing ``` Each line is a discrete action stated in plain terms. When the Agent hits something ambiguous β€” like an unexpected date format β€” it flags the item for review rather than guessing silently. agent reasoning during harmonization ## View summary of harmonization When the Agent finishes, it returns a summary of how columns were mapped: | Category | Columns | | ---------------------------------- | ------- | | Deterministic β€” historical | 66 | | Deterministic β€” direct passthrough | 10 | | AI-mapped (non-NULL) | 12 | | NULL mappings | 77 | **Deterministic β€” historical** mappings match a pattern the system has mapped before. **Deterministic β€” direct passthrough** mappings are unambiguous one-to-one field matches. **AI-mapped** mappings required inference. **NULL mappings** are target fields with no corresponding source data. The agent also returns **DQ check results**, such as: * 37 checks passed (including 1 fixed) * 0 checks failing * 0 SQL errors Where a check fails and the Agent can resolve it automatically, it does so β€” for example, normalizing `obligor_credit_score_type` from `'FICO Score 8 Auto'` to `'FICO Auto'`. Finally, the Agent summarizes **key mappings applied**: how source fields were mapped to the standardized target schema, grouped by category (Identifiers, Dates, etc.). * **Source** β€” the original field name from the input data * **Target** β€” the corresponding field name in the harmonized schema Where a mapping includes a format change (such as reformatting a date) the Agent displays the transformation alongside the field names, so you can verify both the mapping and the conversion in one place. Once harmonization completes, review each mapping before accepting the tape. This step is foundational β€” a poorly reviewed tape undermines every downstream analysis built on it. The goal is to get every mapping to an approved status. ## View review panel The right panel shows mapping review status. Click any mapping to see: * The transformation applied. * An AI explanation of the confidence level. * Any DQ tests passed. * AI Memory β€” previously approved mappings inform future mapping accuracy. The more mappings you approve, the more deterministic future mappings on similar tapes become Mapping detail panel with Source Column picker and Sample Values You can change a mapping by clicking the **Source Column/Expression** drop-down menu, which offers the following options: * **No value (null)** β€” leave the target column unmapped. * **Function Expression** β€” write a custom expression manually * **Use AI to generate expression** β€” have AI generate a transformation instead of a direct field mapping * **Columns** β€” map directly from an available source field A **Sample Values** panel shows real data from the selected source column, so you can confirm the mapping is correct before accepting it. ### Review mapping results Review the overall breakdown, for example: * 62 deterministic (historical) * 15 deterministic (direct pass-through) * 17 AI-mapped * 78 null * 72 high-confidence Click **View More** for the full count table. For each mapping, **Accept**, **override**, or **reject** it, then use **Accept & Next** to move through the list. ### Filter and sort mappings Use the Filter icon (top right) to isolate mappings that need review. You can filter by: * **AI confidence** β€” High, Medium, Low, None, Overridden * **Data quality** β€” Passed, Failed, No DQ checks * **Error status** β€” Has error, No error You can sort by: * Default order * Name (A β†’ Z) * Name (Z β†’ A) * Confidence (low β†’ high) β€” surface what needs the most review * Data quality (worst first) * Status (unmapped β†’ accepted) Filter mappings Sorting by confidence (low β†’ high) or data quality (worst first) prioritizes manual verification on the mappings most likely to need it. ## Correct a mapping via chat To fix an incorrect mapping, describe the issue and the correct value directly in chat β€” for example, correcting a `deal_id` mapping. ``` Updated `deal_id` to `'TAOT_2026_1'`. DQ passed. ``` The agent overwrites the mapping; you then accept it. ## Meta-questions during review You can ask the Agent questions about the review state at any point β€” for example, "how many source columns were there in total and how many were mapped?" ## Add DQ tests The agent can suggest data quality tests to add. Adding these improves the Agent's mapping performance on future tapes, not just the current one. ## Add a new CDM To add a new CDM: 1. Click the **+** at the top of the review panel, then select **CDMs**. 2. In the dialog, enter a name for the common data model and click **Create**. 3. In the new CDM, choose **Add Table** or **Upload Schema**. 4. If you choose **Add Table**, enter names, types, descriptions, and DQ checks for each column. ## Complete the tape 1. Continue reviewing remaining confidence tiers (medium, low, etc.). If you're certain of a mapping, you can specify it directly and check sample values in the window. 2. Once every mapping is accepted, click **Output Preview** to preview the harmonized result. 3. Once satisfied, click **Done**. 4. The tape is marked ready. Download it as Excel or CSV, or add it as a table to your SQL Warehouse. Continue to [Stratifications & Pool Selection](/data-analysis/getting-started/structured-finance/stratifications-pool-selection) to run analyses on the harmonized tape. # Monitoring Source: https://docs.prophecy.ai/data-analysis/production/monitoring Review your deployed projects, scheduled pipelines, and run history The **Observability** interface in Prophecy lets you monitor your deployed projects, review scheduled pipelines, and audit run history. You'll be able to view all projects and pipelines that your teams own. To access, click **Observability** (the line graph icon) in the left sidebar. ## Deployed Projects The **Deployed Projects** tab in the Observability interface shows all projects that have been [published](/data-analysis/development/versioning/version-control) to Prophecy fabrics. It lets you view details on release versions and recent execution status at the project level. For each project deployment, Prophecy provides the following information. | Column | Description | | ------------------- | --------------------------------------------------------------------------------- | | **Project** | Name of the deployed project. Click the name to navigate to the project metadata. | | **Fabric** | Fabric associated with the project deployment. | | **Release Version** | Version number of the deployed project release. | | **Published** | Time since the project was last published. | | **Last Run** | Time since the most recent pipeline execution in the project. | | **Last Run Status** | Success or failure of the last run. | Use the following filters to narrow the results: | Field | Description | | ------------------ | --------------------------------------------------------------------- | | **Search project** | Search for a project by name. | | **Fabric** | Filter results by the fabric (such as `AnalystDBX`, `devDatabricks`). | When a project is published to multiple fabrics, Prophecy creates an individual project deployment for each fabric. You will see all of these deployments in the Deployed Projects page. You cannot revert or delete a project deployment. ## Scheduled Pipelines The **Scheduled Pipelines** tab in the Observability interface displays all pipelines configured with a schedule across environments. It provides details on pipeline execution frequency and recent run outcomes. Use this page to verify that pipelines are scheduled correctly and to identify patterns of success or failure. For each scheduled pipeline, Prophecy provides the following information. | Column | Description | | --------------- | ----------------------------------------------------------------------------------------------------- | | **Pipeline** | Name of the pipeline. Click to view the pipeline in the project editor. | | **Fabric** | Fabric where the pipeline is deployed. | | **Project** | Project that contains the pipeline. | | **Triggers** | Type of trigger for the schedule. Schedules can include multiple triggers. | | **Last 5 Runs** | Status of the 5 most recent runs. To find the history of all pipeline runs, open the Run History tab. | You may see multiple rows for the same pipeline if the parent project is deployed to multiple fabrics. Use the following filters to narrow the results: | Field | Description | | ------------------- | ----------------------------------------------------------- | | **Search pipeline** | Search by pipeline name. | | **Fabric** | Filter by fabric. | | **Project** | Filter by the project that contains the scheduled pipeline. | ## Run History Use the **Run History** tab in the Observability interface to view and analyze recent pipeline runs across all fabrics and projects. This page provides detailed metadata for each run, including its type, duration, status, and execution context. You can also use Run History to open **historical runtime logs** and inspect previous pipeline executions in Historical Mode. Use the following filters to narrow the results. | Field | Description | | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Date Range** | Select a date range to limit results to runs that started in that window. | | **Fabric** | Filter by fabric used to run the pipeline. | | **Project** | Filter by the name of the project that contains the pipeline. | | **Run Type** | Select one or more run types to include. Supported values include:
  • Scheduled
  • Pipeline Run
  • App Run
  • Scheduled App Run
  • API Pipeline Run
| For each historical pipeline run, Prophecy provides the following information. | Column | Description | | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **ID** | Hover the **#** sign to view the full unique pipeline run ID.
You can reference this value in the [Trigger Pipeline API](/api-reference/pipelines/trigger-pipeline-run). | | **Fabric** | Name of the fabric used to execute the pipeline run. | | **Pipeline** | Name of the pipeline that was executed. | | **Project** | Name of the project containing the pipeline. | | **Run Type** | How the run was triggered.
This includes interactive runs, pipeline schedules, APIs, and analysis runs.
Any scheduled runs include additional information about the schedule that triggered the run. | | **Start Time** | Timestamp when the run started. | | **End Time** | Timestamp when the run completed. | | **Duration** | Elapsed time of the run. | | **Result** | Status of the run: `Success` (green) or `Failed` (red). | | **See logs** | Opens historical runtime logs and the associated pipeline snapshot in Historical Mode. | ## Open historical runtime logs You can inspect completed pipeline executions directly from the Run History table. To open historical runtime logs: 1. Open the **Run History** tab and locate the run you want to inspect. 2. Click **See logs** in the far-right column. Open historical runtime logs Prophecy opens the selected pipeline in **Historical Mode**, restoring: * Runtime logs for the selected execution * The historical pipeline snapshot * Component execution statuses and progress information In Historical Mode, the pipeline canvas becomes read-only so you can safely review the historical execution context without modifying the current pipeline version. Historical runtime logs and pipeline snapshots are retained for a limited period of time and may not be available for every historical run. For more information about runtime logs, Historical Mode, and pipeline snapshots, see [Runtime logs](/data-analysis/development/runs/runtime-logs). # Publication concepts Source: https://docs.prophecy.ai/data-analysis/production/publication/publication-concepts Understand drafts, published versions, deployments, and fabrics in SQL and Model projects Publication lets you create versioned releases of a project and deploy those releases to development, staging, or production environments. Publication applies only to SQL projects that use Simple Version Control. Projects that use standard Git workflows manage releases through Git instead. ## Publication workflow Projects move through three lifecycle stages: | Stage | Purpose | | ----------------- | ---------------------------------------------------------------------------- | | Save a draft | Save work in progress | | Publish a version | Create a reusable project release | | Deploy | Make that release available in an environment, such as Dev, Staging, or Prod | ```mermaid theme={null} flowchart LR A[Draft] --> B[Published Version 1.1] B --> C[Dev Deployment] B --> D[Staging Deployment] B --> E[Prod Deployment] ``` Publishing does not automatically run pipelines. After deployment, you can [schedule pipelines](/data-analysis/production/scheduling/scheduling) independently for each fabric. Schedules are associated with deployments, not published versions. ## Save a draft As you develop and edit your project, you can **save drafts** to store in-progress project changes before publication. Use drafts to: * Save work while continuing development. * Collaborate with other users. * Review changes before publication. * Prepare a version for deployment. Saving a draft does not create a deployable project version and does not make changes available in deployment environments. ## Draft pipelines When you mark a pipeline as draft, Prophecy excludes its in-progress changes from the project's next published version β€” it doesn't remove the pipeline itself. For example, if a project includes pipelines `a`, `b`, and `c`, and you mark `b` as draft, the next published version includes your latest changes to `a` and `c`, while `b` remains unchanged in that version unchanged, frozen at whatever was last published for it. If `b` had never been published before you marked it as draft, it's simply left out of every published version until you remove the `draft` designation. "Save a draft" and "draft pipelines" both use the word *draft* but mean different things by the term. Saving a draft records in-progress changes to the whole project. Marking a pipeline as draft excludes that pipeline from publication. Use a draft pipeline to: * Continue developing a pipeline while excluding in-progress changes from the next release. * Publish other pipelines in the project without waiting on unfinished work. While a pipeline is marked as draft: * Any in-progress changes are not included in the next published version. If it's already been published before, it ships unchanged in the new version; if it hasn't, it's left out of the version entirely. * Any app built on that pipeline is held back the same way. * The pipeline can't be renamed or deleted while it's a draft. Publishing the project doesn't remove the pipeline or interrupt it. Its last published version carries into the new release unchanged, and any existing schedule on it keeps running. That is, Prophecy redeploys that schedule's job with the same, unchanged definition as part of the release. When you remove draft status, the pipeline is included in the next version you publish; you don't need to take any special steps for the pipeline to "catch up." To mark or unmark a pipeline as draft, see [Mark Pipeline as Draft in the topic Pipelines for Data Analysis](/data-analysis/development/pipelines/data-analysis-pipelines#mark-pipeline-as-draft). ## Publish a version Publishing creates a versioned release of your project. A published release can be deployed to one or more environments, but publishing alone does not deploy it. Until deployment occurs, the published version remains available in Prophecy but is not yet available in a development or production environment. When you publish a version, Prophecy: * Assigns a version number and description. * Packages the project. * Makes the version available for deployment. Published versions are immutable. After a version is published, it cannot be modified. Any changes require publishing a new version. This allows Prophecy to preserve published versions as stable references for deployment and rollback. This helps you: * Track project changes over time. * Deploy consistent versions across environments. * Roll back to earlier project versions if necessary. ## Deploy a project When you **deploy** a project,you make a published version available for execution in a specific deployment environment. In Prophecy, deployment environments are called [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics). Examples include development, staging, and production environments for platforms such as Databricks, Snowflake, or BigQuery. Different environments often require different configurations, such as connections, schemas, or runtime parameters. You can use [project parameter sets](/data-analysis/development/parameters/parameter-sets) to apply environment-specific configuration during deployment. During deployment, Prophecy: * Builds the project in the selected environment. * Applies the selected project parameter set. * Creates or updates the deployment for that environment. You can deploy different versions of the same project to different environments. For example, a development environment might run version `1.1` while production continues to run version `1.0` until validation is complete. Each environment (such as Dev or Prod) can contain only one deployed version of a project at a time. After deployment, you can [schedule pipelines](/data-analysis/production/scheduling/scheduling) independently for each environment. # Publish and deploy projects Source: https://docs.prophecy.ai/data-analysis/production/publication/publish-project Publish SQL projects and deploy them to fabrics Use publication controls in Prophecy Studio to save drafts, create versioned releases, and deploy those releases to fabrics. For an overview of drafts, published versions, deployments, and fabrics, see [Publication concepts](/data-analysis/production/publication/publication-concepts). Publishing consists of two independent actions: * **Release**: Create a versioned project release. * **Deployment**: Deploy that release to one or more fabrics. You can release a project without deploying it, or release and deploy in the same action. You can also choose the [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics) where you will deploy the project: ```mermaid theme={null} flowchart LR A[Save Draft] --> B[Publish Release] B --> C{Select fabrics?} C -->|No| D[Release Only] C -->|Yes| E[Deploy to Fabric A] C -->|Yes| F[Deploy to Fabric B] ``` A release represents a versioned snapshot of your project. Deployments determine where that release runs. ## Publication controls The publication controls in the SQL IDE change based on the current project state. ### Button labels | Button label | Meaning | | ------------- | --------------------------------------------------------------------------- | | Save to Draft | The project contains unsaved changes that must be saved before publication. | | Publish | The project is ready to publish. | | Share | Professional Edition sharing flow. | Prophecy hides publication controls when: * The project is a template or example project * The project is in shared project mode * The project is in replay mode * The project is in historical mode * Simple Version Control is not enabled for the project Under certain conditions, publication controls are disabled. | Condition | Tooltip | | ------------------------------ | --------------------------------------------- | | Viewing version history | "Publish is disabled in version history view" | | Code generation in progress | "Please wait for code generation to finish" | | Project is read-only or locked | "Publish is disabled in read-only view" | | Agent is in read-only mode | No tooltip | ## Save draft changes If your project contains unsaved changes, the publication control displays **Save to Draft**. To save draft changes: 1. Click **Save to Draft**. 2. Review the draft changes. 3. Save the draft. Saving a draft stores the latest project changes without publishing or deploying the project. ## Publish and deploy a project After you save your latest draft changes, the publication control displays **Publish**. To publish a project version: 1. Click **Publish**. 2. Review the version details. 3. Review the changes included in the version. 4. Enter or edit the version number. 5. Add a version description. 6. Optional: Select one or more fabrics for deployment. 7. Optional: Select a project parameter set. 8. Click **Publish**. If any pipelines are marked as draft, their in-progress changes aren't included in this version β€” they ship unchanged (or are left out entirely, if never published before). The changes list shows them with a **Draft - Held** label, along with any business apps built on them, and a summary line reports how many pipelines were held back. To include a pipeline's latest changes in this publish, remove its draft status first. See [Draft pipelines](/data-analysis/production/publication/publication-concepts#draft-pipelines). Fabric selection determines whether or not the project is deployed. If you publish without selecting a fabric, Prophecy creates the published version but does not deploy it to an environment. When you select one or more fabrics during publication, Prophecy creates or updates deployments for those fabrics using the published version. You can only deploy to fabrics your team can access. To deploy a project to a fabric, your team must have access to that fabric. For example: * Development fabrics may allow broader access for testing and iteration * Production fabrics may restrict deployment access to approved users If your production fabric uses Databricks connections, consider using a service principal for authentication. This helps scheduled pipelines run reliably in production environments. During deployment, Prophecy: * Builds the project in the selected fabric. * Applies the selected project parameter set. * Creates or updates the deployment for that environment. Each deployment represents a specific project version running in a specific fabric. The same project version can be deployed to multiple fabrics, and each deployment is tracked independently. After deployment, you can [schedule pipelines](/data-analysis/production/scheduling/scheduling) independently for each fabric. Pipeline schedules are environment-specific. When you deploy a project to multiple fabrics, each fabric can have its own schedule configuration. Prophecy validates the project before publication. | Validation issue | Message | | ---------------------- | --------------------------------------------------------- | | Compilation errors | "Cannot publish: Please fix all compilation errors first" | | Invalid version number | "Please enter a valid version number" | | Missing description | "Please add a description to publish the version" | | No fabric selected | "Please select at least one fabric to publish" | | No deployment changes | "Please update fabric(s) or parameter set to publish" | Prophecy generates logs for each publication step. If deployment fails, publication logs show the exact step where the process stopped. The logs help you troubleshoot failed publications and identify where publication stopped. The publication process includes these steps: 1. **Fetching fabric info**: Retrieve information about the selected fabric 2. **Reading project code**: Review project code elements 3. **Packaging project**: Bundle project components together 4. **Connecting to deployment service**: Connect to the deployment service 5. **Deploying to fabric**: Deploy the project to the selected fabric Publish logs ## Request to publish Prophecy allows multiple users to work on the same project simultaneously. In collaborative projects, publishing often requires coordination with other users. If another user currently holds a peer lock on the project, Prophecy opens a **Request to Publish** dialog instead of publishing immediately. Collaborators can review, approve, or reject the publication request before the new version is released. ## Share a project Available for [Free and Professional Editions](/data-analysis/administration/platform/editions) only. Professional Edition users can share SQL projects with: * Existing Prophecy users. * External users who are not yet registered with Prophecy. Shared projects open in a read-only view. External users can: * review the project, * interact with Agent in limited mode, * and sign up for Prophecy. Users cannot modify shared projects directly. To make changes, users must clone the project after signing in to Prophecy. To share a project: 1. Click the arrow next to **Save to Draft** or **Publish**. 2. Select **Share project**. 3. Share the project by: * entering an email address, * or copying the share link. Guest users can ask a limited number of Agent questions in shared projects. Guest Agent chats are not persisted and disappear after the browser refreshes. # Email alerts Source: https://docs.prophecy.ai/data-analysis/production/scheduling/alerts Report the outcome of a scheduled pipeline run The built-in Prophecy scheduler supports email alerts for pipeline schedules. You can configure alerts to notify you or others when a scheduled pipeline run starts, succeeds, or fails. ## Parameters When you enable **Alerts on the full job** in a schedule, configure the alert with the following parameters. | Parameter | Description | | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Email | One or more email addresses that Prophecy will sent the alert to. | | Trigger alerts on | When the alert will run. You can select multiple triggers.
  • **Start**: Send an email when the pipeline run begins
  • **Success**: Send an email if the pipeline run succeeds
  • **Failure**: Send an email if the pipeline run fails
| # Pipeline gem Source: https://docs.prophecy.ai/data-analysis/production/scheduling/pipeline-trigger-gem Start pipeline runs from a gem in the canvas The Pipeline gem allows you to run another pipeline from within your current pipeline. It supports conditional triggering, parameter passing, and returning metadata for each pipeline run. This gem is useful for building orchestrated workflows directly in the visual canvas. You can find it under the **Custom** category in the gem drawer. A single Pipeline gem execution can trigger the target pipeline multiple times. Triggered pipeline runs execute sequentially, and the Pipeline gem completes only after all triggered runs finish. You can create pipelines solely dedicated to pipeline orchestration using multiple instances of this gem. We recommend labeling your orchestration pipelines to differentiate them from standard data pipelines. ## Limitations The Pipeline gem has the following limitations: * It can only trigger pipelines for the current project. * The [Monitoring](/data-analysis/production/monitoring) page doesn't show the parent pipeline that triggered a run using the Pipeline gem. ## Input and Output The following table describes what the Pipeline gem expects as input and what it will produce as output. | Port | Description | | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **in0** | Optional input dataset that may include a `status` column and additional columns.

Valid statuses include `success`, `failure`, or `skipped`.

Each row of the input triggers a separate pipeline run per Pipeline gem execution. | | **out** | One output dataset that contains the metadata for each triggered pipeline run.

The Pipeline gem returns one row per triggered run. | When no input is connected: * The Pipeline gem triggers the child pipeline once. * The output contains a single row of metadata for that run. When an input dataset is connected: * If a `status` column is present, it is used to evaluate [trigger conditions](#trigger-conditions). If the trigger condition is met, the Pipeline gem triggers the child pipeline once for each input row. * If a `status` column is **not** present, the Pipeline gem always triggers the child pipeline once for each input row. * Additional columns can be used to pass parameter values into each run. * The output contains one row of metadata per input row. The input dataset can be any dataset (including the output of an upstream Pipeline gem). ### Output schema Each row in the output table represents a pipeline run triggered from the Pipeline gem. The output schema provides you with the following information for each pipeline run. | Column | Type | Description | | --------------- | ------------ | --------------------------------------------------------------------------- | | `status` | String | Final status of the triggered pipeline: `success`, `failure`, or `skipped`. | | `pipelineRunID` | String | Unique ID of the triggered pipeline run. | | `startTime` | Timestamp | Time when the triggered pipeline started. | | `endTime` | Timestamp | Time when the triggered pipeline finished. | | `error` | String | Error message, if any. | | `logs` | String Array | Log messages from the triggered pipeline. | Pipelines are `skipped` when the input does not meet the trigger condition set in the gem. ## Parameters The Pipeline gem accepts the following parameters. | Parameter | Description | | ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Pipeline to run | Select a pipeline to trigger in the list of pipelines from the same project. | | Version | Choose which published version of the target pipeline to trigger. See [Version](#version) below. | | Trigger only if | Choose a [condition](#trigger-conditions) to control when the trigger fires based on the `status` column of the input dataset (if present). If the condition is not met, the Pipeline gem runs, but no pipelines are triggered. The output will show that all pipeline runs were `skipped`. | | Maximum number of pipeline triggers | Set a maximum number of times the child pipeline runs per Pipeline gem execution. Maximum is `10,000`. | | Set pipeline parameters | Set values for the [pipeline parameters](/data-analysis/development/parameters/parameters) defined in the child pipeline. You can specify constants, expressions, or column values for each parameter. If you don't specify a value, the child pipeline uses its default parameter value. | Use multiple input rows to launch runs of the same pipeline with different sets of parameter values. ### Version By default, the Pipeline gem triggers whatever version of the target pipeline is currently in the working tree. You can choose to lock it to a specific published version using the **Last Released** option. | Option | Description | | ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Current** (default) | Triggers the target pipeline's current working-tree state. Always up to date, including unpublished changes. | | **Last Released** | Resolves to the target pipeline's newest published version at the moment each run executes. | | A specific version (e.g. `3`) | Locks the trigger to that exact published release. The Pipeline gem keeps triggering that version even after the target pipeline is edited, drafted, or republished again β€” until you change this setting. | Changing **Pipeline to run** resets **Version** back to **Current**. Only published versions that include the target pipeline are selectable. Versions that don't include it appear disabled in the list. ### Trigger conditions A Pipeline gem requires at least one non-empty input row to execute. If input is empty, the pipeline will not run, even when trigger condition is set to **Always run**. | Trigger condition | Description | | --------------------------- | ---------------------------------------------------------------------------------------------- | | **Always run** (default) | Executes the pipeline whenever at least one input row is present, regardless of input status. | | **All pipelines succeeded** | Executes only when every input row has a `success` status. | | **All pipelines failed** | Executes only when every input row has a `failure` status. | | **All pipelines finished** | Executes after all input pipeline runs have completed, regardless of their success or failure. | | **Any pipeline succeeded** | Executes if at least one input row has a `success` status. | | **Any pipeline failed** | Executes if at least one input row has a `failure` status. | ## Execution behavior Prophecy supports the following execution behaviors for Pipeline gems. | Execution behavior | Description | | --------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Sequential execution | If a Pipeline gem triggers the same target pipeline multiple times, the triggered pipeline runs execute sequentially. The Pipeline gem completes only after all triggered runs finish. | | Sequential triggering | Pipeline gems execute in sequence when they are connected sequentially or assigned sequential [gem phases](/data-analysis/gems/#gem-phase). Downstream Pipeline gems wait for upstream Pipeline gems to finish before triggering. | | Parallel triggering | Pipeline gems can execute in parallel when separate pipeline branches use the same gem phase. | | Recursive triggers | Pipeline gems support recursive execution. If a triggered pipeline contains another Pipeline gem, that Pipeline gem executes normally. | # Schedule setup Source: https://docs.prophecy.ai/data-analysis/production/scheduling/schedule-setup Follow the steps to configure the Prophecy scheduler Use Prophecy's built-in scheduler to automate pipeline runs based on time or file triggers. You can create multiple schedules for a single pipeline, each with its own configuration including parameter sets, triggers, and alerts. ## Prerequisites To leverage the built-in scheduler, your project must use the [Simple Git Storage Model](/data-analysis/development/versioning/git-storage-model). Prophecy's built-in scheduler does not appear in projects that use the Normal Git Storage Model. ## Configure a schedule To configure a schedule for a pipeline: 1. Open a pipeline in the project editor. 2. Expand the **...** menu in the pipeline canvas header. 3. Click **Schedule**. 4. To create a new schedule, click **New Schedule**. To edit an existing schedule, select it from the schedules list. 5. Under **Which version runs**, choose whether the schedule always uses the latest published version or locks to one specific version. See [Lock a schedule to a version](#lock-a-schedule-to-a-version) below. 6. Expand the **Choose Pipeline Parameter Set** dropdown and select the pipeline [parameter set](/data-analysis/development/parameters/parameter-sets) you want to use for this schedule. 7. Configure the [trigger](/data-analysis/production/scheduling/triggers) for the schedule. 8. Optionally, set up an [email alert](/data-analysis/production/scheduling/alerts) in your schedule. 9. Click the **On** toggle to enable the schedule. 10. Click **Schedule** to save the schedule. You can create multiple schedules for the same pipeline, each with different parameter sets, triggers, and timing configurations. This allows you to run the same pipeline with different settings at different times or frequencies. ## Lock a schedule to a version By default, a schedule always runs whichever version of the project is currently published. That is, when the team publishes a new version, the schedule automatically moves to this version for the next run. You can instead *lock* a schedule to one specific published version, so the scheduler keeps running that version no matter how many times the project is republished. To change or remove a lock, open the schedule and update **Which version runs**: * Select **Always the latest published** to unlock the schedule and have it follow future publishes again. * Select **Lock to a specific version** and choose a different version to move the lock forward. A locked schedule that's already running a locked version is unaffected by publishing. Publishing makes no change to it until you update the lock yourself. Locks apply per fabric. The same schedule can be locked to a version on one fabric's deployment while following the latest published version on another. ## Activate schedules Enabling a schedule does **not** automatically activate it. To activate your schedules, [publish the project](/data-analysis/production/publication). When you publish a project, Prophecy deploys all enabled schedules to your selected fabrics. In other words, your scheduled pipelines will begin running in the fabrics you selected. Any time you make changes to a schedule (including enabling, disabling, or editing the configuration), you must republish the project for those changes to take effect. ## Delete a schedule To delete a schedule: 1. Open a pipeline in the project editor. 2. Expand the **...** menu in the pipeline canvas header. 3. Click **Schedule**. 4. Select the schedule you want to delete. 5. Click **Delete Schedule**. 6. Click **Delete** to confirm the deletion. When you delete a schedule, you still have to publish the project for the changes to take effect in the deployed pipeline. ## What's next Once you have set up your schedules: * Monitor automated pipeline runs in the [Observability](/data-analysis/production/monitoring) interface. * Set up [email alerts](/data-analysis/production/scheduling/alerts) to stay informed about your scheduled pipeline runs. * Learn more about different [trigger types](/data-analysis/production/scheduling/triggers). # Scheduling Source: https://docs.prophecy.ai/data-analysis/production/scheduling/scheduling Automate your pipeline runs using schedules Prophecy supports the following ways to automate pipeline execution. * **Built-in scheduler**: Use [Prophecy Automate](/data-analysis/administration/platform/architecture) to orchestrate pipelines from the UI. * **Pipeline gem**: Configure the [Pipeline gem](/data-analysis/production/scheduling/pipeline-trigger-gem) to start pipeline runs from a gem. * **API**: Call the [Trigger Pipeline API](/api-reference/pipelines/trigger-pipeline-run) to start pipelines from external systems. This page describes how to use the **built-in scheduler** in SQL projects. ## Overview In Prophecy, a schedule automates the execution of a single pipeline in a project. Scheduled pipelines run in [fabrics](/data-analysis/environment/fabrics/prophecy-fabrics) (execution environments) that are defined during project publication. Each schedule defines: * When the pipeline should run. Prophecy supports time-based or file-based [triggers](/data-analysis/production/scheduling/triggers). * Optional [email alerts](/data-analysis/production/scheduling/alerts) to report the outcome of the pipeline run. To schedule multiple pipelines, create a separate schedule for each one. You can enable or disable schedules individually, but enabling a schedule doesn't activate it. An enabled schedule becomes active **only when the parent project is published**. Similarly, disabling a schedule also requires republishing the project, since schedule status is part of the deployment configuration. ## Schedule activation Enabling a schedule in your project doesn't immediately activate it. For a schedule to take effect, you must first [publish the project](/data-analysis/production/publication). This is because publishing defines how and where scheduled pipelines are deployed and executed. When you publish a project, you do two key things: * Select one or more fabrics. These are the environments where scheduled pipelines will run. A separate deployment is created for each fabric β€” publishing to one fabric does not affect other deployments. If you do not select any fabrics during project publication, **no deployments will be created**. As a result, no scheduled executions will occur, even if a schedule has been configured. * Specify the project version. This version will be deployed to the fabric. You can either create a new version or deploy a previously published version. Scheduling flow ## Monitor scheduled pipelines You and your team members might have many scheduled pipelines in your Prophecy environment. The [Observability](/data-analysis/production/monitoring) interface in Prophecy includes the following information: * List of deployed projects * List of pipeline schedules per fabric * History of pipeline runs and run status You'll only see information about projects owned by your teams. ## Authentication lifespan Scheduled pipelines run without human intervention, which makes them vulnerable to failures caused by expired or invalid credentials. If Prophecy cannot authenticate a connection used by the pipeline due to an error such as an expired token or deleted user, the pipeline run will fail. This risk increases in complex environments where: * Pipelines depend on multiple external data sources. * The same schedule is deployed across multiple fabrics. * Fabrics store different credentials for the same connection. To ensure reliable scheduled runs, only deploy to fabrics that use connection credentials that won't expire unexpectedly. For Databricks connections, [consider using a service principal](/data-analysis/environment/connections/databricks#authentication-methods) for authentication. Service principals are designed for authorizing access to Databricks resources when running unattended processes. ## What's next To learn more about using the Prophecy-native scheduler, explore the following pages. * [Set up schedule](/data-analysis/production/scheduling/schedule-setup) * [Schedule trigger types](/data-analysis/production/scheduling/triggers) * [Email alerts](/data-analysis/production/scheduling/alerts) * [Pipeline gem](/data-analysis/production/scheduling/pipeline-trigger-gem) # Trigger types Source: https://docs.prophecy.ai/data-analysis/production/scheduling/triggers Learn about different trigger types for schedules Prophecy's built-in scheduler allows you to orchestrate your pipelines with user-defined automation logic. Schedules support two trigger types: * [Time-based](#time-based-trigger): The scheduled pipeline runs at defined intervals. * [File arrival or change](#file-arrival-or-change-trigger): The scheduled pipeline runs only if a file arrives or changes. ## Time-based trigger ### Frequency Time-based trigger configurations vary according to the frequency of the trigger. The following sections describe the parameters needed for each frequency. When using time-based triggers, the default timezone is the timezone from where you access Prophecy. #### Minute | Parameter | Description | Default | | ------------ | ------------------------------------------------------------------------------------- | -------- | | Repeat every | The interval in minutes between pipeline runs.
Example: Repeat every 10 minutes. | 1 minute | #### Hourly | Parameter | Description | Default | | --------------------- | --------------------------------------------------------------------------------------------------------------------------- | -------------------------------------- | | Repeat every ... from | The interval in hours between pipeline runs, starting at a specific time.
Example: Repeat every 2 hours from 12:00 AM. | Every 1 hour
starting at 12:00 AM | #### Daily | Parameter | Description | Default | | --------- | --------------------------------------------------------------------------- | ------- | | Repeat at | The time of day when the schedule will run.
Example: Repeat at 9:00 AM | 2:00 AM | #### Weekly | Parameter | Description | Default | | --------- | ---------------------------------------------------------------------------------------------------- | -------- | | Repeat on | The day(s) of the week that the pipeline will run.
Example: Repeat on Monday, Wednesday, Friday | Sunday | | Repeat at | The time of the day that the pipeline will run.
Example: Repeat at 9:00 AM | 12:00 AM | #### Monthly | Parameter | Description | Default | | --------- | ------------------------------------------------------------------------------------------------ | -------- | | Repeat on | The day of the month that the pipeline will run.
Example: Repeat on the first of the month. | 1 | | Repeat at | The time of the day that the pipeline will run.
Example: Repeat at 9:00 AM | 12:00 AM | #### Yearly | Parameter | Description | Default | | ------------- | ---------------------------------------------------------------------------------------------------- | -------- | | Repeat every | The day and month that the pipeline will run each year.
Example: Repeat every March 15. | None | | Repeat on the | The specific occurrence of a day in a given month.
Example: Repeat on the third Monday of June. | None | | Repeat at | The time of the day that the pipeline will run.
Example: Repeat at 9:00 AM | 12:00 AM | ## File arrival or change trigger Use this trigger type to run a scheduled pipeline only when a new file is added or an existing file changes in a specified directory. Prophecy supports file change detection (data change sensors) for [S3](/data-analysis/environment/connections/s3) and [SFTP](/data-analysis/environment/connections/sftp) connections only. ### Trigger configuration Each trigger configuration defines where and how Prophecy detects file changes. Provide the following: | Parameter | Description | | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Configuration name | A label to help you identify this trigger configuration. | | Connection | The **S3** or **SFTP** connection where the system should watch for file changes. | | File path | The full path to the directory being monitored. The pipeline runs when a file is added or modified in this location. You can use either a fixed directory path, such as `/user/documents/`, or a regular expression to match specific file patterns.

Example: `/user/*/*.{csv,jar}` finds all subdirectories under` /user` and watches files ending in `.csv` or `.jar`. | #### Add multiple trigger configurations You can define multiple trigger configurations to monitor more than one location for file updates. When you add more than one configuration, the pipeline uses `AND` logic, whereby all conditions must be met before the pipeline runs. For example, you may want to add the following trigger configurations: | Configuration name | Connection | File path | | ------------------ | ---------- | ----------------------------- | | Transactions | S3 | `/finance/transactions/*.csv` | | Exchange Rates | SFTP | `/finance/fx_rates/` | With this setup, the pipeline runs only when a change is detected for both Transactions and Exchange Rates. ### Advanced settings These advanced settings determine how the poll-based data sensors work. For additional details, jump to the following sections. | Parameter | Description | | ------------------------- | ------------------------------------------------------------------------------------------- | | Poke Interval (seconds) | How often to check the condition. Defaults to 60 seconds. | | Sensor Timeout (seconds) | Maximum duration for the sensor to keep checking. If not set, the sensor runs indefinitely. | | Exponential Backoff Retry | Enables increasing delay between polling attempts to reduce system load. | | Poll Mode | Determines whether to use **Poke** or **Reschedule** mode. | #### Poke Interval The Poke Interval specifies how frequently the sensor checks for file updates, such as new files appearing or existing files changing. The default interval is 60 seconds. #### Sensor Timeout The Sensor Timeout value determines the maximum amount of time the sensor continues to check for file updates. By default, the sensor runs indefinitely. Set a timeout only if there's a point at which you know the check is no longer useful or relevant. #### Exponential Backoff Retry Sometimes, the sensor does not detect any file updates when it checks the directory. When Exponential Backoff Retry is enabled, the wait time between checks increases every time the sensor must retry the operation after failure (no file updated). When active, the delay between each polling attempt increases exponentiallyβ€”e.g., 60 seconds, 2 minutes, 4 minutes, 8 minutesβ€”while adding randomness (jitter) to avoid system-wide contention. This option is enabled by default to reduce system load. #### Poll Mode Choose between the following modes: * **Poke**: Keeps resources active between checks. Best for short intervals when quick response times are needed. * **Reschedule**: Frees resources and reschedules itself after each interval. Recommended for long intervals to reduce resource usage. # Data diff Source: https://docs.prophecy.ai/data-engineering/ci-cd/data-diff View the difference between a target dataset and an expected dataset Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Data diffs can help you identify when pipeline outputs do not meet predefined expectations. This can be useful for: * Understanding the success of your pipeline migrations. * Catching discrepancies before deploying pipelines to production. * Detecting unintended schema changes in datasets over time. * Identifying data mismatches when troubleshooting data transformation issues. ## Prerequisites To compute a data diff, you need: * **Prophecy 4.0 or later.** * **A Prophecy Python project.** Data diffs are not supported for Scala or SQL projects. * **ProphecyLibsPython 1.9.37 or later** as a dependency for your project. For more information, see [Prophecy libraries](/data-engineering/extensibility/dependencies/prophecy-libs). * **An expected dataset**. It must be in Parquet or Databricks Catalog Table format. Data diff cannot be computed on [Databricks Serverless](/data-engineering/fabrics/spark-provider/databricks/databricks-serverless). ## What is data diff? Data diffs are outputs of Target gems that show you differences between your target table and an expected table. Similar to a data sample, you can explore the data diff after [interactively running](/data-engineering/development/runs/execution) a pipeline. The data diff has four views: * **Overview**: A summary that displays various high-level comparisons of the generated (target) and expected datasets. Review the following section to understand each statistic in more detail. * **Column differences**: The dataset schemas and the number of matching values for each column in both datasets. * **Values differences**: A table that displays side-by-side differences of every value in both datasets. This will show a sample of the data. * **Data samples**: Samples of the generated and expected datasets for data exploration. Data diffs only temporarily appear in the pipeline. They are not persisted in your project. ### Overview The following table provide in-depth descriptions of each statistic in the Overview tab of the data diff. | Field | Description | | -------------------------- | ---------------------------------------------------------------------------------------------------------------------- | | Datasets matching status | Whether the datasets are matching or not. | | Number of columns matching | The number of columns that match. Two columns match if they have the same name and the same set of values. | | Number of rows matching | The number of rows that match. Two rows match if they have the same key column(s) and the same values for each column. | | Location | The location of the generated and expected datasets. | | Primary keys | The primary keys defined in the data diff configuration. | | Number of columns | The number of columns in the generated and expected datasets. | | Number of rows | The number of rows in the generated and expected datasets. | | Unique primary keys | The number of unique primary keys in the generated and expected datasets. | | Duplicate primary keys | The number of duplicate keys in the generated and expected datasets. | ### Unique and duplicated primary keys Prophecy only calculates the data diff on rows with unique primary keys. Why is that? Assume you have a the following table, where `first_name` and \*last\_name\` are the primary keys: | `first_name` | `last_name` | `cust_id` | | ------------ | ----------- | --------- | | John | Smith | 23542 | | John | Smith | 49203 | | Jane | Doe | 43291 | | Jane | Brown | 09312 | If you try to compute the data diff **John Smith**, how will you know which row is the correct match? It is impossible to match the rows with 100% confidence. Because of this ambiguity, **Prophecy ignores rows with duplicated primary keys in the data diff.** ## Configuration Data diffs are configured in **Target** gems. 1. Open a Target gem in your pipeline. 2. In the top right of the gem dialog, click the **Options** (ellipses) menu. 3. Select **Data Diff**. This adds the Data Diff step to your Target gem configuration. 4. Open the **Data Diff** step. 5. Fill in the required parameters and **Save** the gem. Review the data diff configuration parameters in the following table. | Parameter | Description | | -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Specify the key columns to join datasets on | The column(s) that will match rows between datasets. In other words, rows are "joined" on these columns for comparison. Rows with duplicated primary keys **will not** be included in the data diff calculation. You can view the number of unique and duplicated primary keys in the **Overview** tab of the data diff. | | Specify an alternative Parquet dataset path | The path to the expected dataset in Parquet format. | | Specify an alternative dataset Catalog Table | The location of the Catalog Table in Databricks. This includes the database, schema (Unity Catalog only), and table names. | The row order of the generated and expected dataset does not matter, as the rows are joined by keys, rather than row order. ## Enable or disable If you have configured data diff in your Target gem, Prophecy will automatically generate the data diff output. However, you can disable this feature from the gem action menu if needed. Disabling data diff (without deleting it) can be useful for large datasets, as it helps reduce computation time. # Deployment workflow Source: https://docs.prophecy.ai/data-engineering/ci-cd/deployment/deploy-project Learn how to use Git for deployment Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Prophecy provides a recommended mechanism for using Git based development. The four main phases of integrating your changes are **Commit**, **Pull**, **Merge**, and **Release**. A standard development pattern looks like this, though other mechanisms like forking are also supported: Project Git Flow Let's develop and deploy a project to illustrate these phases. ## Create a project 1. Open the **Create Entity** page from the left sidebar. 2. Click **Project** 3. Create a new project or import an existing project. 4. Choose existing Git credentials or connect new Git credentials. 5. Specify the desired repository and path accessible to your Git user to store the project. For new projects, specify an empty repository or an empty path within an existing repository. For imported projects, select a repository, forked repository, or repository path that already contains the relevant project code. New project ## Checkout branch A branch in Git is like a separate version of your project where you can make changes without affecting the main version. It lets you work on new features or fixes independently, and you can later merge your changes back into the main project. Let's checkout (or switch to) a different branch. 1. Open a project in the editor. 2. Select the **Git** tab in the bottom bar. 3. Click **Checkout branch**. 4. Choose an existing branch from the dropdown or create a new branch by typing a new name. 5. Click **Checkout**. Checkout branch ## Make changes on the branch A **commit** represents changes to one or more files in your project that let you keep and view project history. When you make changes to the pipeline, you will want to commit them. Let's see how to commit pipeline changes to preserve them in Git. 1. Make a change in your project, such as creating a new pipeline. 2. Click **Commit Changes** to open the Commit window. 3. Write a commit message or let our Copilot write one for you. 4. Review the change and click **Commit**. Commit changes Once you have committed your changes, you have the option to continue developing your pipelines or to **merge** your changes. In this case, choose **Continue**. ## Merge changes **Merge** will take the changes in the *current branch* and merge them into a *base branch*. Your changes will become part of the base branch and will be available to anyone else whose can access the base branch. 1. Ensure that you are merging to the correct base branch. 2. Review the commits that you are merging. 3. If everything looks right, click **Merge**. Merge changes In this how-to, we have not discussed the **Pull** section of the Commit window. **Pull** brings changes that have occurred in [remote branches](https://git-scm.com/book/ms/v2/Git-Basics-Working-with-Remotes) into the Prophecy-local branches. If you have any upstream changes that need to be pulled into the local branches, you'll see that option in the **Pull** section before moving on to **Merge**. ## Release and Deploy When you **Release and Deploy** your project, a particular commit is tagged in the base branch with a user-specified version. This allows you designate a new version as ready for production, or inform users who may be subscribed to datasets defined within your project that there might be changes in the published dataset. 1. Select the commit you wish to deploy. 2. Specify the release version. This is usually a number, using a strategy like [semantic versioning](https://semver.org/). 3. Fill in the release notes to describe the release. 4. Click **Release and Deploy**. Release changes At this point, you have worked through one iteration of your project's lifecycle! To learn more about different deployment options, visit [Deployment](/data-engineering/ci-cd/deployment/deployment). # Project release and deployment Source: https://docs.prophecy.ai/data-engineering/ci-cd/deployment/deployment Release projects and deploy jobs Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Once you have developed and tested your custom components like gems, pipelines, models, or jobs in Prophecy, the next step is to make them available for use. This involves Releasing and Deploying them to the respective environments. You can Release and Deploy via the Prophecy UI or you can use the [Prophecy Build Tool](/data-engineering/ci-cd/prophecy-build-tool/prophecy-build-tool) CLI to integrate with other CI-CD tools. Let's see how you can do it via the Prophecy UI below. ## Requirements You must be a [team admin](/data-engineering/administration/management/teams/teams) to release and deploy a project. ## Overview What happens when you click **Release and Deploy** in the [Git workflow](/data-engineering/ci-cd/git/git) of your project? ### Release The release step ensures that your new project code is synced to your Git repository. * A [Git tag](https://git-scm.com/book/en/v2/Git-Basics-Tagging) is created with the version you specify. * The new project version is pushed to your Git repository. ### Deploy The deploy step builds all the pipelines and gems, uploads the artifacts, and schedules all the jobs in the project to the respective environments. * Pipelines are compiled and built into an artifact (Wheel file for Python or Jar file for Scala). These artifacts are then uploaded to your environment. * Gems (including custom gems) are built and uploaded to an internal artifactory. They aren't directly copied to your environments, as they are used in generating code for the pipelines, not during job/pipeline execution. However, the code for gems does get committed to your Git repo as part of the project. * Depending on the type of job, jobs are copied to their respective environments as JSON files for Databricks jobs and as Python DAGs for Airflow. There are no specific deployment steps needed for other project entities. ## Advanced Settings For most users, a regular project release takes care of both the release and deployment of pipelines, gems, and jobs to respective environments. However, if you want more control over the deployment process, you can edit project settings. 1. Open the **Settings** page in the project metadata. 2. Open the **Deployment** subtab. 3. Change the deployment mode or enable/disable unit testing. Deployment settings ### Deployment modes Deployment modes dictate how the release and deploy steps will be controlled. * **Deploy on Release (default)**: Release and deploy in a single step. * **Staged Release and Deployment**: Separate release and deploy into two separate steps. * **Selective Job Deployment**: Select specific jobs during the Deploy step. Use this if you have many jobs in a project and only want to deploy a few at a time. *Only the deployed jobs will use the latest versions of pipelines, datasets, and subgraphs.* ### Unit tests Writing good [unit tests](/data-engineering/ci-cd/tests) is key for data pipeline quality and management. When you enable unit tests for deployment, unit tests will run as part of pipeline builds. This might lead to a slight increase in the build time. ## History You can view the release and deployment history in the **Releases & Deployments** tab of your project metadata. Releases and Deployments The page includes the following subtabs. * **Releases**: Find a history of releases including information about the author, creation time, and latest tag. You can also view the logs of the latest deployment associated with that tag. * **Current Version**: View the current state of all deployed jobs per environment. Select the fabric to view the list of all jobs deployed in that environment, along with their versions and deployment logs. * **Deployment History**: See the history of all past deployments, along with the time that it was deployed and related logs. ## What's next Follow the tutorial [Develop and deploy a project](/data-engineering/ci-cd/deployment/deploy-project) to try to deploy a project yourself! # External release tags Source: https://docs.prophecy.ai/data-engineering/ci-cd/deployment/use-external-release-tags Use external release tags for deployment and dependency in Prophecy Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. If you use external CI-CD tools to merge and release your projects, you can use [release tags](https://git-scm.com/book/en/v2/Git-Basics-Tagging) from those tools within Prophecy for deployment and dependencies. Once you've deployed an external tag in Prophecy, you can add that release as a dependency in a project. ## View external release tags Any externally created release tag that you pull into Prophecy is visible on the **Releases & Deployments** tab of your project metadata. External_tags_list Tags that are created externally are labeled with an **(1) External** tag. If your latest tags aren't showing, click on the **(2) Refresh** button to refresh the list of tags. ## Deploy an external release tag To deploy an existing tag, follow these steps: 1. From `...` in the top right corner, select **(1) Deploy**. This opens the Deploy dialog. Deploy button 2. Select a release version you wish to deploy by using the **(1) Choose a release** dropdown. Once you select a version, the table below shows the jobs that are going to be modified (there might not be any jobs). Click **(2) Deploy** to start the deployment. Deploy start If you have enabled [Selective Job Deployment](./deployment#deployment-modes), then you can pick the jobs you wish to deploy. Additionally, you have the option to override the fabrics for these jobs. Job selection **is not required** to deploy the release. This deploys a new release. You can access deployment logs from the Deployment History tab. ## Use the release as a dependency Once you've deployed a tag, you can use the release as a dependency in a project. First, navigate to the **(1) Dependencies** tab of the relevant project and click **(2) + Add Dependency**. Add dependency Next, in the **Create Dependency** dialog: 1. For **Type**, select **Package Hub Dependency**. 2. For **Name**, choose the project that contains the external release tag. 3. For **Version**, select the option that matches your external release tag. 4. Click **Create Dependency**. Add another dependency You can also edit a dependency and update its version to an externally released version. ## FAQ **How does Prophecy support tags from a repo that is linked to multiple Prophecy projects?** A Git tag is a pointer to a specific commit in the repo. It's not linked to a subfolder in the repo. So in this case, if you create a tag, it would be available for all projects linked to the repo. **Do the tags have to follow a certain pattern to be recognized?** No. Prophecy supports all tag patterns supported in Git. Prophecy automatically recognizes external tags after you visit the Release and Deployment page or refresh the page. # Git best practices Source: https://docs.prophecy.ai/data-engineering/ci-cd/git/best-practices Learn about what we recommend to do if you are working with Git. ## Overview To minimize and ideally avoid conflicts in your pull requests: 1. Have each user create their own feature branches. 2. Keep feature branches modular, which means: * Only modify assets as needed on your feature branch * Avoid modifying assets that other users are modifying in parallel feature branches ## Deep dive: Branching Strategy Use a proper Git branching strategy that aligns with your team's goals. ### Small teams Small team branching strategy For small teams: 1. Have a main branch for developers to merge in their changes. 2. (Optional) Create a release branch to use as a staging area for releases across different environments. 3. Check out a new feature branch off the main branch for any new features. 4. Pull in the latest changes and resolve all merge conflicts on your feature branch before merging it into the main branch. 5. Have one developer work on a pipeline at any given time. A Prophecy merge conflict resolution occurs at the entity level, and a pipeline is considered as a single entity. 6. Have your admin enable pull requests for specific projects. This encourages teams to perform code reviews and collaborate, which ensures higher code quality and fewer bugs. ### Large teams Large team branching strategy For large teams with rigid and separate execution environments, and promotion processes: 1. Create the following branches: `main`, `release`, `develop`, and feature branches 2. Correspond each branch to different execution environment: For example, the following branches may correspond to the following execution environments: | Branch | Execution environment | | ----------------------- | ---------------------- | | `main`, or `hotfix` | `prod` | | `release` | `uat`, `test`, or `qa` | | `develop`, or `feature` | `dev` | # Set up Git credentials Source: https://docs.prophecy.ai/data-engineering/ci-cd/git/git Connect external repositories to Prophecy Prophecy utilizes [Git](https://git-scm.com/book/en/v2/Getting-Started-About-Version-Control) to align with DevOps practices. Git lets you: * Store your visual pipelines as code. * Track your project metadata, including workflows, schedules, datasets, and computed metadata. * Collaborate on your projects and perform code reviews. * Track changes across time. ## Projects and Git repositories When you create a project in Prophecy, you must choose an empty Git repository or a path in a repository to host the underlying project code. This way, Prophecy can keep track of any changes to the project (including pipelines, models, datasets, and jobs) in a systematic way. If you're unfamiliar with Git, or you don't need to connect to an external Git repository, you can connect your projects to **Prophecy-managed** repositories. To leverage the full functionality of Git, however, you need to connect your projects to external Git repositories. To do so, you need to add external Git credentials to Prophecy. ## Add Git credentials to Prophecy When you create a project in Prophecy, You can either choose a Prophecy-managed repository or connect an external repository. To add an external Git account to your Prophecy environment: 1. Open **Settings** from the ellipses menu in the left sidebar. 2. Select the **Git** tab. You will see any existing Git credentials here. 3. Click **Add new**. 4. Choose the relevant **Git Provider** and provide the following details: * An alias for your Git credentials. * Your Git email. * Your Git username. * Your Git personal access token (PAT). You should be able to find instructions on accessing your personal access token in each external provider's documentation. 5. Click **Connect** to save the credentials. You can also enter new Git credentials directly when creating a new project. ### Share credentials In Prophecy, you can share your Git credentials with teams so they can access Git repositories without signing in to your Git provider (for example, GitHub or GitLab). This allows them to work with Prophecy projects in external repositories while keeping your Git provider account private. Team members with shared access can: * Connect new Prophecy projects to repositories available to the connected Git account. * Access existing Prophecy projects already connected to repositories in the connected Git account. * Migrate Prophecy projects from Prophecy-managed Git to an external repository in the connected Git account. In **Settings > Git**, you'll see your credential split into three sections: * **Personal Git Credentials**: Credentials you own that are available only to you.. * **Team Credentials (Owned by you)**: Credentials you own and have shared with teams. * **Team Credentials (Shared with you)**: Credentials owned by others and shared with you. To share credentials, select the **Share credentials with teams** checkbox when adding or editing your Git credentials, and choose the teams you want to share them with. You can only edit the credentials you own. You cannot modify credentials that have been shared with you. ### GitHub Oauth If you choose GitHub as your external Git provider, you can add your Git credentials using GitHub Oauth. To use GitHub Oauth, a GitHub admin will need to [authorize Prophecy](https://docs.github.com/en/apps/oauth-apps/using-oauth-apps/authorizing-oauth-apps) to access the APIs required to access your organization's repositories. Follow the [approval steps](https://docs.github.com/en/organizations/managing-oauth-access-to-your-organizations-data/approving-oauth-apps-for-your-organization) to set this up. ## Fork per User When you create a project, you have the option to choose a single repository shared among users, or to **Fork per User**. When you Fork per User, every user gets their own [fork](https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/fork-a-repo) of the repository. When you fork a repository, you must already have both the upstream repository and a Fork per User repository present. * Changes made in forked repository do not affect the upstream repository. * Please follow the normal Git flow for raising pull requests to the original repository from the forked repository. ## What's next Learn more about the [Git workflow](/data-engineering/ci-cd/git/git-workflow) or try it yourself in [Develop and deploy a project](/data-engineering/ci-cd/deployment/deployment). # Resolve conflicts Source: https://docs.prophecy.ai/data-engineering/ci-cd/git/git-resolve Resolve conflicts that you may run into while merging your changes This page describes how to resolve conflicts that you may run into while merging your changes. There are two ways to resolve merge conflicts when they arise: * Merge using the Prophecy interface * Merge in your external Git interface ## Merge in Prophecy There are a couple of ways to manually resolve merge conflicts in Prophecy. ### Left or right merge strategy The Left or Right merge strategy gives you a the option to resolve the conflict by choosing one version of your code to keep. After choosing, you can click **Next** to continue the merge process. Choose a Git conflict manual merge * **(A)** **Strategy**: You must choose a preferred strategy to resolve the conflict. Here the Left strategy keeps the version on branch `master`, while the Right strategy keeps the version on branch `dev`. * **(B)** **Open on master**: Clicking this opens the model on branch `master` for you to view. * **(C)** **Open on dev**: Clicking this opens the model on branch `dev` for you to view. Here are the read-only views on branch `master` on the left and branch `dev` on the right: View Git conflict merge strategies ### Code changes merge strategy For SQL, you can also toggle on **Code Changes** to view and edit the code before validating. You can resolve conflicts by making code changes directly on the files. View Git conflict merge strategies Once you've made the changes that you want to keep, click **Next**. The merge process will compile the files. Errors caused by conflict resolution In rare cases, your merge attempt may result in an error after the compile completes. You'll be asked to fix the error before proceeding. See **Diagnostics** at the bottom for details on what the error is and how you might go about fixing it. Once you've fixed the error, click **Try Again**. If you're confident that the errors are fine to leave as is, click **Ignore Errors**. ## Remote and local branch conflicts If your commit history differs between a remote and local branch (due to actions like rebasing and reverting), one way to resolve this is to delete one of the branches. To delete the local branch: 1. Open the **Git** menu from the project footer. 2. Choose **Delete branch**. 3. Select the branch you wish to delete. 4. Click the **Delete local branch** button. 5. Click **Delete** to confirm deletion. Delete a Git branch To delete the remote branch, do so from your external Git provider. # Git workflow Source: https://docs.prophecy.ai/data-engineering/ci-cd/git/git-workflow Follow the Git workflow in your Prophecy project You can interact with a project's Git workflow from the project metadata page or within the project editor. Git workflow ## Checkout In Prophecy, you cannot make edits directly on the `main` branch of your project. Instead, you have to make changes on a development branch and merge those changes into the main branch. Therefore, the Git workflow begins by creating and checking out a new branch. Git checkout When you are selecting a branch to checkout, you might be able to select branches that have been created remotely. Once you checkout a remote branch, it will be cloned locally, and it will no longer show up in the list of remote branches. ## Commit When you make changes to your pipeline, you need to commit these changes to save them. You can view these changes either visually, or using the **Code changes** view. View Git changes | **Feature** | **View** | **Description** | | -------------------- | -------- | -------------------------------------------------------------------------------------------- | | Branch history | Visual | Shows previous commits on the branch. | | Entities changed | Visual | Explains which entities were modified. | | Code changes toggle | Code | Allows you to view the code differences, with highlighted lines for additions and deletions. | | Code changes tab | Code | Displays all of the files with changes. | | Metadata changes tab | Code | Displays all of the Prophecy metadata files with changes. | | Reset | Both | Reverts the changes. | | Commit Message | Both | Explains the changes that will be saved in this commit. | ## Pull Sometimes, you will be able to go straight from committing your changes to merging your changes. However, there are a few steps you might need to complete before merging your changes: 1. Pull **remote** changes into the **local** current branch. 2. Pull **remote** changes into the **local** base branch. Note that the base branch will be `main` by default. 3. Pull changes from the **local** base branch into the **local** current branch. You will not be able to complete this step before pulling remote commits. Git pull Once you complete these steps, you might run into merge conflicts. If that happens, you can [use the Prophecy interface](/data-engineering/ci-cd/git/git-resolve) to resolve them. Before you pull remote changes into local branches, you will have to commit (or discard) your local changes. ## Merge Once you have committed your changes, you have the ability to **merge** them to a different branch. If you merged your branch in your external repository, you can tell Prophecy that you did so. Merge branch ### Pull Requests If you need code reviews or use protected branches like main, you can enable pull requests during the merge step. This lets you open a pull request in your external Git provider directly from Prophecy. To enable and configure this workflow, see [Pull requests](/data-engineering/ci-cd/git/pull-request-templates). ## Release and Deploy Once the changes are merged, you can [release and deploy](/data-engineering/ci-cd/deployment/deployment) a branch from the Prophecy user interface. ## Rollback or restore If you have changes that you do not want to commit, there are a few ways to discard them. 1. Click the Reset button in the commit stage of the Git workflow. This discards all changes in your project since the last commit. 2. Restore a particular component from the Project Browser. 3. Rollback changes to a particular component in the Git workflow. 4. Rollback to a specific commit in the Git workflow. Rollback component # Pull requests Source: https://docs.prophecy.ai/data-engineering/ci-cd/git/pull-request-templates Open Pull Requests from within Prophecy By default, Prophecy lets you merge changes from a development branch directly into the base branch. This is fast and works well for smaller teams or rapid iteration. However, in larger teams or production environments, you may want more control over how code is reviewed, approved, and integrated. Enabling pull requests allows you to: * Review and approve changes before they are merged * Enforce branch protection rules on repositories like GitHub or Bitbucket When pull requests are enabled for a project, Prophecy generates a merge URL based on your configured template. This lets you open external pull requests directly within Prophecy projects. This page describes how to enable and use pull requests in more detail. Pull request support is only available for projects connected to an external Git provider. It's not supported for Prophecy-managed Git. ## Enable pull request template To use pull requests in a project, you need to enable pull request templates for that project. 1. Open your project metadata. 2. Open the **Settings** tab. 3. Next to **Pull Request Template**, toggle on the **Enabled** button. 4. Review the template URL. PR template settings The PR template URL requires two variables which are used to build a URL string. The `{{source}}` variable represents the active development branch, and the `{{destination}}` variable represents the base branch to which the development branches need to be merged to, like `main`. ### Template examples Using this template: ```shell theme={null} https://github.com/exampleOrg/exampleRepo/compare/{{destination}}...{{source}}?expand=1 ``` An example pull request URL generated from the above template for merging a branch named `feature` to branch `main` would look like: ```shell theme={null} https://github.com/exampleOrg/exampleRepo/compare/main...feature?expand=1 ``` Using this template: ```shell theme={null} https://bitbucket.org/exampleOrg/exampleRepo/pull-requests/new?source={{source}}/1&dest={{destination}} ``` An example pull request URL generated from the above template for merging a branch named `feature` to branch `main` would look like: ```shell theme={null} https://bitbucket.org/exampleOrg/exampleRepo/pull-requests/new?source=feature/1&dest=main ``` ## Open pull request in Prophecy After you have enabled pull requests for a project, you will see the option to create pull requests directly in Prophecy. When you open the Git dialog of a project: * The **Merge** step will be replaced by an **Open Pull Request** step. * The **Open Pull Request** button on this screen will open an external pull request in a new tab. If you run into issues, ensure that your PR [template](#enable-pull-request-template) is configured correctly. Once you merge the branch remotely in the pull request, you need to let Prophecy know that this step is complete. 1. Return to the **Open Pull Request** step of the Git dialog. 2. Click **Merged Externally**. 3. Click **Confirm**. Merged externally ## Set Version Before Merge When your base branch (such as `main`) is protected, direct commits are not allowed. This can interfere with release processes that rely on version bumps made directly on the main branch after merging. To support these workflows, Prophecy now allows you to set the next version **before** merging your development branch. When **Pull Request Template** is enabled for a project: 1. In the **Git** dialog, go to the **Open Pull Request** step. 2. Select the **Incremental Project Version** checkbox. 3. Select a version or type a new version. This sets the version of your development branch. 4. Open the pull request and merge your changes. The base branch now has the correct version. Once merged, Prophecy will auto-fill the release version on the **Release** screen of the Git dialog based on the version you set. You can directly proceed to release and no additional changes will be committed to the base branch. # PBT on GitHub Actions Source: https://docs.prophecy.ai/data-engineering/ci-cd/prophecy-build-tool/pbt-github-actions Example usage of Prophecy Build Tool on GitHub Actions Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. ## Using PBT with GitHub Actions Prophecy Build Tool (PBT) can be integrated with GitHub Actions to: * validate pipelines * build artifacts (`.jar` / `.whl`) * run unit tests * deploy pipelines to Databricks View an [Example GitHub repository](https://github.com/prophecy-samples/external-cicd-template). ## Prerequisites * A Prophecy project hosted in a GitHub repository * A Databricks workspace for deployment ## Configuration ### Environment variables PBT requires the following: * `DATABRICKS_HOST` * `DATABRICKS_TOKEN` Store the token as a GitHub Actions secret: > Settings β†’ Secrets β†’ Actions β†’ New repository secret Then reference it in your workflow: ```yaml theme={null} env: DATABRICKS_HOST: "https://" DATABRICKS_TOKEN: ${{ secrets.DATABRICKS_TOKEN }} ``` ## Example workflow (deploy on push to prod) This workflow: * runs on every push to `prod` * validates, builds, and tests pipelines * deploys artifacts to Databricks Create the workflow file: ``` .github/workflows/exampleWorkflow.yml ``` ### Workflow definition ``` name: Example CI/CD with GitHub actions on: push: branches: - "prod" env: DATABRICKS_HOST: "https://sample_databricks_url.cloud.databricks.com" DATABRICKS_TOKEN: ${{ secrets.PROD_DATABRICKS_TOKEN }} FABRIC_ID: "4004" # replace with your fabric id jobs: build: runs-on: ubuntu-latest steps: - uses: actions/checkout@v3 - name: Set up JDK 11 uses: actions/setup-java@v3 with: java-version: "11" distribution: "adopt" - name: Set up Python uses: actions/setup-python@v4 with: python-version: "3.9.13" - name: Install dependencies run: | python3 -m pip install --upgrade pip pip install build pytest wheel pytest-html pyspark==3.3.0 prophecy-build-tool - name: Validate pipelines run: pbt validate --path . - name: Build pipelines run: pbt build --path . - name: Run tests run: pbt test --path . - name: Deploy pipelines run: pbt deploy --path . --release-version 1.0 --project-id example_project_id ``` ## What this workflow does 1. Triggers on pushes to the `prod` branch 2. Sets required environment variables for Databricks access 3. Installs Java, Python, and PBT dependencies 4. Validates pipeline syntax (`pbt validate`) 5. Builds pipelines into `.jar` / `.whl` artifacts (`pbt build`) 6. Runs unit tests (`pbt test`) 7. Deploys artifacts and jobs to Databricks (`pbt deploy`) * Uploads artifacts referenced in `databricks-job.json` * Creates or updates Databricks jobs * Deploys pipeline configurations to DBFS if defined If any step fails, the workflow stops and the run is marked as failed. # PBT on Jenkins Source: https://docs.prophecy.ai/data-engineering/ci-cd/prophecy-build-tool/pbt-jenkins Example Usage of Prophecy Build Tool on Jenkins Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. This example shows how to use Jenkins to: * **validate and test** Prophecy pipelines on pull requests * **deploy pipelines** to Databricks environments after merge Each environment (`develop`, `qa`, `prod`) maps to a separate Databricks workspace. Typical promotion flow: ``` feature β†’ develop β†’ qa β†’ prod ``` ## Prerequisites You should have: * A Git repository containing a Prophecy project * A Jenkins server with permission to create pipelines * Databricks workspaces for each environment ### Required Jenkins plugins * GitHub Pull Request Builder (for test pipeline) * GitHub plugin (for deploy pipeline) > Check plugin compatibility with your Jenkins version before installing. ## Configuration ### Secrets Configure the following credentials in Jenkins: * `DEMO_DATABRICKS_HOST` * `DEMO_DATABRICKS_TOKEN` * `PROD_DATABRICKS_HOST` * `PROD_DATABRICKS_TOKEN` ### Fabric ID Find your Fabric ID from: > Metadata β†’ Fabrics β†’ \ ## Testing pipeline (PR validation) This pipeline: * runs on pull requests to `develop`, `qa`, and `prod` * validates pipelines * runs unit tests ### Trigger Use **GitHub Pull Request Builder** to trigger on: * new PRs * updates to PRs ### Jenkinsfile (test) ``` // .jenkins/deploy-declarative.groovy pipeline { agent any environment { PROJECT_PATH = "./hello_project" VENV_NAME = ".venv" } stages { stage('checkout') { steps { git branch: '${ghprbSourceBranch}', credentialsId: 'jenkins-cicd-runner-demo', url: 'git@github.com:prophecy-samples/external-cicd-template.git' sh "apt-get install -y python3-venv" } } stage('install pbt') { steps { sh """ python3 -m venv $VENV_NAME source ./$VENV_NAME/bin/activate pip install -U pip build pytest wheel pytest-html pyspark prophecy-build-tool """ } } stage('validate') { steps { sh ". ./$VENV_NAME/bin/activate && python3 -m pbt validate --path $PROJECT_PATH" } } stage('test') { steps { sh ". ./$VENV_NAME/bin/activate && python3 -m pbt test --path $PROJECT_PATH" } } } } ``` ### What this pipeline does 1. Checks out the PR branch. 2. Installs PBT and dependencies. 3. Validates pipeline syntax. 4. Runs unit tests. ## Deploy pipeline (post-merge) This pipeline: * runs on commits to `develop`, `qa`, `prod`. * deploys pipelines to the corresponding Databricks environment. ### Trigger Use a **GitHub webhook** to trigger on push events. ### Jenkinsfile (deploy) ``` // .jenkins/test-declarative.groovy def DEFAULT_FABRIC = "1174" def fabricPerBranch = [ prod: "4004", qa: "4005", develop: DEFAULT_FABRIC ] pipeline { agent any environment { DATABRICKS_HOST = credentials("${env.GIT_BRANCH == "prod" ? "DEMO_PROD_DATABRICKS_HOST" : "DEMO_DATABRICKS_HOST"}") DATABRICKS_TOKEN = credentials("${env.GIT_BRANCH == "prod" ? "DEMO_PROD_DATABRICKS_TOKEN" : "DEMO_DATABRICKS_TOKEN"}") PROJECT_PATH = "./hello_project" VENV_NAME = ".venv" FABRIC_ID = fabricPerBranch.getOrDefault("${env.GIT_BRANCH}", DEFAULT_FABRIC) } stages { stage('install pbt') { steps { sh """ python3 -m venv $VENV_NAME source ./$VENV_NAME/bin/activate pip install -U pip build pytest wheel pytest-html pyspark prophecy-build-tool """ } } stage('deploy') { steps { sh ". ./$VENV_NAME/bin/activate && python3 -m pbt deploy --fabric-ids $FABRIC_ID --path $PROJECT_PATH" } } } } ``` ### What this pipeline does 1. Selects the target environment based on branch. 2. Installs PBT. 3. Builds pipelines into `.jar` / `.whl` artifacts. 4. Uploads artifacts to Databricks. 5. Creates or updates jobs. ## Notes * Each `sh` step runs in a separate shell, so the virtual environment must be reactivated. * For Scala pipelines, ensure **JDK 11** is installed on Jenkins nodes. * Jenkins files are stored in the repository; Jenkins stores only triggers and credentials. # Use Prophecy Automate with PBT Source: https://docs.prophecy.ai/data-engineering/ci-cd/prophecy-build-tool/pbt-prophecy-automate Use Prophecy Build tool with Prophecy Automate Some projects may be able to use [Prophecy Automate](/data-engineering/administration/platform/architecture#what-is-prophecy-automate) with the [Prophecy Build Tool](/data-engineering/ci-cd/prophecy-build-tool/prophecy-build-tool). Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Limited availability. For more information, reach out to our [Sales team](mailto:sales@prophecy.io). The following matrix outlines supported environments for using the `prophecy-automate` runtime with the Prophecy Build Tool: | Capability / Requirement | Linux x86\_64 | Linux aarch64 (Graviton) | | ------------------------ | ----------------- | ------------------------ | | Wheel available | βœ” | βœ” | | Minimum `glibc` | 2.28 (Rocky 8) | 2.28 (Rocky 8) | | Minimum Linux kernel | 3.17+ | 3.17+ | | Python 3.10 | βœ” (DBR 14) | βœ” (DBR 14) | | Python 3.11 | βœ” (DBR 15) | βœ” (DBR 15) | | Python 3.12 | βœ” (DBR 16) | βœ” (DBR 16) | | Spark 3.3–3.5 | βœ” | βœ” | | Tableau target support | βœ” | βœ– | | Build method | Docker (Rocky 8) | Docker (cross-compile) | | Spark 4.x | Not yet supported | Not yet supported | # Prophecy Build Tool (PBT) Source: https://docs.prophecy.ai/data-engineering/ci-cd/prophecy-build-tool/prophecy-build-tool Prophecy Build tool Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. The **Prophecy Build Tool (PBT)** is a command-line utility for building, testing, validating, and deploying Prophecy-generated projects. The Prophecy Build Tool lets you integrate Prophecy pipelines into existing CI/CD systems (such as GitHub Actions or Jenkins) and orchestration platforms (such as Databricks Workflows). You can use the PBT to run the same set of tasks described in [Project release and deployment](/data-engineering/ci-cd/deployment/deployment) from the command line or in a CI/CD script. The Prophecy Build Tool only supports building **Python** and **Scala** projects. ## Features Using the Prophecy Build tool, you can: * Build pipelines (all or a subset) in Prophecy projects. * Unit test pipelines in Prophecy projects. * Deploy jobs with built pipelines on Databricks. * Deploy jobs filtered by fabric IDs on Databricks. * Integrate with CI/CD tools like GitHub Actions. * Verify the project structure of Prophecy projects. * Deploy pipeline configurations. * Add Git tags to a deployment. * Set versions for PySpark projects. ## Requirements To install and run the Prophecy Build Tool, you need: * `pip` * `python >=3.7` (Recommended 3.9.13) * `pyspark` (Recommended 3.3.0) ### Using generated wheel artifacts The Prophecy Build Tool (PBT) generates Python wheel (`.whl`) files as part of the build process. These artifacts must be available in the execution environment where your pipeline runs. Depending on your setup, you may need to install or provide the wheel manually. #### Local or ad-hoc execution If you are running jobs locally or outside a managed deployment pipeline, install the wheel file before execution: ``` pip install dist/.whl ``` Alternatively, when using Spark: ``` spark-submit --py-files dist/.whl ... ``` #### Databricks or cluster environments Ensure the wheel is either: * attached as a cluster library, or * included in your deployment process. ## Install PBT To install the Prophecy Build Tool, run: ``` pip3 install prophecy-build-tool ``` See the [PyPI package](https://pypi.org/project/prophecy-build-tool/) for the latest version. ## Usage ```shell theme={null} Usage: pbt [OPTIONS] COMMAND [ARGS]... Options: --help Show help. Commands: build Build Prophecy pipelines build-v2 Same as build but with additional options deploy Deploy pipelines and jobs deploy-v2 Same as deploy but with additional options test Run unit tests validate Validate pipelines for diagnostics versioning Add versions to PySpark pipelines tag Create a Git tag for the version in `pbt_project.yml` ``` ## Configuration Before using PBT with Databricks, set the following environment variables: ```shell theme={null} export DATABRICKS_HOST="https://example_databricks_host.cloud.databricks.com" export DATABRICKS_TOKEN="exampledatabrickstoken" ``` These variables define the Databricks workspace and credentials used for building and deploying. ## Build pipelines The `build` command compiles all or specific pipelines within a Prophecy project. ```shell theme={null} pbt build --path /path/to/your/prophecy_project/ ``` You can also use a v2 version of the command, which adds an `--add-pom-python` option. ```shell theme={null} pbt build-v2 --path /path/to/your/prophecy_project/ ``` ### Build options | Option | Description | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------- | | `--path TEXT` | **Required.** Path to the directory containing the `pbt_project.yml` file. | | `--pipelines TEXT` | Comma-separated list of pipelines to build. | | `--ignore-build-errors` | Continue even if build errors occur.
Refer to logs for details. | | `--ignore-parse-errors` | Continue even if pipeline parsing errors occur.
Returns success (`EXIT_CODE = 0`).
Refer to logs for details. | | `--add-pom-python` | Available with `--build-v2`. Adds `pom.xml` and `MAVEN_COORDINATES` files to PySpark builds. | | `--help` | Show help for this command. | To build only specific pipelines: ```shell theme={null} pbt build --pipelines customers_orders,join_agg_sort --path /path/to/your/prophecy_project/ ``` To continue despite build or parsing errors: ```shell theme={null} pbt build --path /path/to/your/prophecy_project/ --ignore-build-errors --ignore-parse-errors ``` If any pipeline fails to build, the Build tool exits with code 1 unless error-skipping flags are used. ## Deploy pipelines and jobs The `deploy` command builds and deploys Prophecy pipelines and jobs to your Databricks workspace. ```bash theme={null} pbt deploy --path /path/to/your/prophecy_project/ --release-version 1.0 --project-id 10 ``` PBT supports the `--release-version` and `--project-id` parameters, used to replace placeholders in your job definition file (`databricks-job.json`). These values determine the DBFS path where artifacts are uploaded. Use the project's ID (from its URL) and a unique release version for each deployment. ### Sample deploy output ```shell theme={null} Prophecy Build Tool v1.0.4.1 Found 1 job: daily Found 1 pipeline: customers_orders (python) Building 1 pipeline 🚰 Building pipeline pipelines/customers_orders [1/1] βœ… Build complete! Deploying 1 job ⏱ Deploying job jobs/daily [1/1] Uploading customers_orders-1.0-py3-none-any.whl to dbfs:/FileStore/prophecy/artifacts/... Updating existing job: daily βœ… Deployment completed successfully! ``` ### Deploy dependent projects Use `--dependent-projects-path` to include dependent Prophecy projects located in subdirectories. ```bash theme={null} pbt deploy --path /path/to/your/prophecy_project/ --release-version 1.0 --project-id 10 --dependent-projects-path /path/to/dependent/prophecy/projects ``` ### Deploy by fabric ID Use `--fabric-ids` to deploy jobs associated with specific Fabric IDs (helpful for multi-workspace environments). ```bash theme={null} pbt deploy --fabric-ids 647,1527 --path /path/to/your/prophecy_project/ ``` You can find Fabric IDs in the Prophecy UI by visiting the Metadata page of a Fabric and checking its URL. ### Skip builds To deploy previously built pipelines without rebuilding: ```bash theme={null} pbt deploy --skip-builds --path /path/to/your/prophecy_project/ ``` ### Deploy specific jobs By default, all jobs are deployed. To deploy selected jobs, use `--job-ids`. ```bash theme={null} pbt deploy --path /path/to/your/prophecy_project/ --job-ids "TestJob1,TestJob2" ``` The Prophecy Build Tool automatically identifies and builds only the pipelines required by those jobs. You can also use a v2 version of the command, which adds several options described in the table below. ```bash theme={null} pbt deploy-v2 --path /path/to/your/prophecy_project/ --job-ids "TestJob1,TestJob2" ``` ### Deploy options summary | Option | Description | | -------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | `--path TEXT` | **Required.** Path containing the `pbt_project.yml` file. | | `--dependent-projects-path TEXT` | Path containing dependent Prophecy projects. | | `--release-version TEXT` | Release version tag for this deployment. | | `--project-id TEXT` | In `deploy`, Prophecy project ID (used to replace placeholders). In `deploy-v2`, path to the directory containing the `pbt_project.yml` file. | | `--prophecy-url TEXT` | Prophecy base URL for deployment. Removed in `build-v2`. | | `--fabric-ids TEXT` | Comma-separated Fabric IDs to filter jobs. | | `--skip-builds` | Skip building pipelines. | | `--job-ids TEXT` | Comma-separated list of Job IDs to deploy. | | `--conf-dir TEXT` | Available with `--deploy-v2`. Path to configuration file folders. | | `--release-tag TEXT` | Available with `--deploy-v2`. Specify a release. tag. | | `--skip-pipeline-deploy` | Available with `--deploy-v2`. Skip pipeline deployment and deploy only job definitions. | | `--migrate` | Available with `--deploy-v2`. Migrates a v1 project to v2.format. | | `--artifactory TEXT` | Available with `--deploy-v2`. Allows use of PyPI/Maven packages instead of DBFS files for deployment. | | `--skip-artifactory-upload` | Available with `--deploy-v2`. Skips uploading to private artifactory (must be used with `--artifactory`). | | `--help` | Show help for this command. | ## Test pipelines The Prophecy Build Tool supports unit testing of pipelines within a Prophecy project. Tests run with the default configuration under `configs/resources/config`. ```bash theme={null} pbt test --path /path/to/your/prophecy_project/ ``` ### Test options | Option | Description | | ---------------------------- | -------------------------------------------------------------- | | `--path TEXT` | **Required.** Path containing the `pbt_project.yml` file. | | `--driver-library-path TEXT` | Path to JARs for `prophecy-python-libs` or other dependencies. | | `--pipelines TEXT` | Comma-separated list of pipelines to test. | | `--help` | Show help for this command. | If `--driver-library-path` is omitted, dependencies are fetched automatically from Maven Central. ### Sample test output ```shell theme={null} Prophecy Build Tool v1.0.1 Found 1 job: daily Found 1 pipeline: customers_orders (python) Unit Testing pipeline pipelines/customers_orders [1/1] ============================= test session starts ============================== platform darwin -- Python 3.8.9, pytest-7.1.2 collected 1 item test/TestSuite.py::CleanupTest::test_unit_test_0 PASSED [100%] ============================== 1 passed in 17.4s =============================== βœ… Unit test for pipeline: customers_orders succeeded. ``` ## Validate pipelines Validation checks all pipelines in a project for warnings and errors, similar to Prophecy’s in-IDE diagnostics. This helps ensure pipelines are production-ready before deployment. ```bash theme={null} pbt validate --path /path/to/your/prophecy_project/ ``` ### Validate options | Option | Description | | ---------------------------- | --------------------------------------------------------- | | `--path TEXT` | **Required.** Path containing the `pbt_project.yml` file. | | `--treat-warnings-as-errors` | Treat warnings as errors during validation. | ## Applying versions to PySpark projects PySpark projects often rely on specific versions of Spark, Python libraries, and data connectors. Versioning the project (via `pbt_project.yml` or `setup.py`) ensures compatibility between your code and these dependencies, helping avoid runtime errors when pipelines are deployed to different environments (local, Databricks). The Prophecy Build Tool lets you set various options for versioning as follows. ### Versioning options | Option | Description | | -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `--path ` | **Required.** Path to the directory containing the `pbt_project.yml` file. | | `--repo-path ` | Path to the repository root. If left blank, the tool uses `--path`. | | `--bump [major\|minor\|patch\|build\|prerelease]` | Bumps one of the semantic version numbers for the project and all pipelines based on the current value. Only works if existing versions follow [Semantic Versioning](https://semver.org/). | | `--set TEXT` | Explicitly set the exact version. | | `--force`, `--spike` | Bypass errors if the version set is lower than the base branch. | | `--sync` | Ensure all files are set to the same version defined in `pbt_project.yml`. *(Implies `--force`.)* | | `--set-suffix TEXT` | Set a suffix string (e.g., `-SNAPSHOT` or `-rc.4`). If this is not a valid semVer string, an error will be thrown. | | `--check-sync` | Check to see if versions are synced. Exit code `0` = success, `1` = failure. | | `--compare-to-target`, `--compare ` | Checks if the current branch has a greater version number than the `` provided. Returns `0` (true) or `1` (false). Also performs a `--sync` check.
**Note:** If `--bump` is also provided, it compares versions and applies the bump strategy if the current version is lower. | | `--make-unique` | Makes a version unique for feature branches by adding build-metadata and prerelease identifiers.
*Format:* `MAJOR.MINOR.PATCH-PRERELEASE+BUILDMETADATA`
*Examples:*
Python β†’ `3.3.0 β†’ 3.3.0-dev0+sha.j0239ruf0ew`
Scala β†’ `3.3.0 β†’ 3.3.0-SNAPSHOT+sha.j0239ruf0ew` | | `--pbt-only` | Apply version operation to `pbt_project.yml` file only. Applicable with `--compare`, `--make-unique`, `--bump`, `--set`, or `--set-suffix`. | | `--help` | Show help for this command. | ## Tagging builds The `pbt tag` command creates a Git tag for the version listed in `pbt_project.yml`. This tag marks a specific point in your project’s history so you can track or redeploy that version later. By default, the tag name includes the branch (for example, `main/1.4.0`) and is pushed to the remote automatically. You can change or remove the branch name with `--branch`, create a custom tag with `--custom`, or skip pushing with `--no-push`. ### Tag options | Option | Description | | ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | | `--path TEXT` | Path to the directory containing the `pbt_project.yml` file **\[required]** | | `--repo-path TEXT` | Path to the repository root. If left blank, it will use `--path`. | | `--no-push` | By default, the tag will be pushed to the origin after it is created. Use this flag to skip pushing the tag. | | `--branch TEXT` | Normally, the tag is prefixed with the branch name: `/`. This option overrides ``. Provide `""` to omit the branch name. | | `--custom TEXT` | Explicitly set the exact tag using a string. Ignores other options. | | `--help` | Show help for this command. | ## Sample output ```shell theme={null} Prophecy Build Tool v1.0.3.4 Project name: HelloWorld Found 1 job: default_schedule Found 4 pipelines: customers_orders, report_top_customers, join_agg_sort, farmers-markets-irs Validating 4 pipelines Validating pipeline pipelines/customers_orders [1/4] Pipeline validated: customers_orders ... βœ… All pipelines validated successfully. ``` ## Quick reference | Command | Description | Common Options | | --------------------- | ------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | `pbt build` | Build all or selected pipelines within a Prophecy project. | `--path` (required)
`--pipelines` *(comma-separated)*
`--ignore-build-errors`
`--ignore-parse-errors` | | `pbt test` | Run unit tests for pipelines using the default configuration. | `--path` (required)
`--pipelines` *(comma-separated)*
`--driver-library-path` *(optional)* | | `pbt validate` | Validate pipelines for warnings or errors before deployment. | `--path` (required)
`--treat-warnings-as-errors` | | `pbt deploy` | Build and deploy pipelines and jobs to Databricks. | `--path` (required)
`--release-version`
`--project-id`
`--fabric-ids`
`--job-ids`
`--skip-builds` | | pbt version | Set versions for PySpark projects. | | | Environment variables | Required for Databricks connections. | `DATABRICKS_HOST`
`DATABRICKS_TOKEN` | ### Example workflow ```bash theme={null} # 1. Build project pipelines pbt build --path /path/to/project # 2. Run tests pbt test --path /path/to/project # 3. Validate before deployment pbt validate --path /path/to/project # 4. Deploy pbt deploy --path /path/to/project --release-version 1.0 --project-id 123 ``` ## What's next To learn how to integrate the Prophecy Build Tool with GitHub Actions and Jenkins, see the following examples: * [GitHub Actions](/data-engineering/ci-cd/prophecy-build-tool/pbt-github-actions) * [Jenkins](/data-engineering/ci-cd/prophecy-build-tool/pbt-jenkins) # CI/CD strategies Source: https://docs.prophecy.ai/data-engineering/ci-cd/reliable-ci-cd Set up CI/CD in Prophecy using environment separation and testing Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Learn how to set up continuous integration and deployment ([CI/CD](https://en.wikipedia.org/wiki/CI/CD)) for your Prophecy data pipelines using environment separation and testing. This page outlines how to: * Configure multi-environment deployments * Set up testing and validation * Deploy pipelines using Prophecy's built-in tools or external CI/CD systems ## Prerequisites To set up CI/CD successfully in Prophecy, ensure that you have: * Access to target execution environments (Databricks workspaces, for example) * Empty Git repositories configured for your projects * Understanding of your data pipeline requirements and SLAs ## Environment setup We recommend setting up separate environments (fabrics) for each stage of deployment. The table below illustrates one example of a multi-fabric setup. | Fabric | Purpose | Data | Access | Cluster Size | | ----------- | --------------------------------------- | -------------------------------- | --------- | ------------ | | Development | Feature development and initial testing | Synthetic or anonymized datasets | Dev team | Small | | QA/Staging | Testing and validation | Production-like data samples | QA team | Medium | | Production | Live data processing | Real production data | Prod team | Large | Data pipeline See Also To learn more about the relationship between fabrics, projects, and teams, visit [Team-based Access](/data-engineering/administration/management/users/access/team-based-access). While you are not required to have multiple execution environments in Prophecy, it is best practice to keep development and production data and access separate. ## Pipeline development workflow Here is an example pipeline development workflow for teams that leverage multiple environments. ### 1. Create a project Start by creating a new project in Prophecy. The project should have its own dedicated Git repository, which helps maintain a clean version history and enables collaboration across teams. The project should be assigned to a team that includes every user who will need to access this project throughout its lifecycle. This is possible because the project team can differ from the fabric team, which should be more restrictive. While you can technically use an empty directory to host your Prophecy project instead, this is not recommended. ### 2. Develop and test Develop pipelines in the development environment using the `dev` branch. This stage involves building pipelines and configuring pipeline parameters to handle different runtime scenarios or environment-specific values. Execute pipelines interactively during development and inspect runtime logs to debug and validate pipeline behavior. During development, you can also create jobs that will automate pipeline execution after deployment. ### 3. Deploy to QA Once development and initial validation are complete, the QA team can begin validating the pipelines by running them interactively in the QA environment. Then, changes to the project can be merged to the `main` branch and deployed in the QA environment to test if scheduled pipelines run as expected. This stage ensures that your pipelines and jobs function as expected in a controlled, production-like setting before they are released to live systems. ### 4. Deploy to production After the QA team has validated the project, the project can be deployed to production. Typically, a small, designated platform team is responsible for this step. The project is deployed to the production fabric, where jobs operate on real production data and run at scale. Jobs in production should ideally execute using a service principal rather than a user identity, since it is an unattended operation. ## Project deployment options You can deploy Prophecy projects using either the built-in Git-based workflow or through an external CI/CD system using the Prophecy Build Tool (PBT). Both approaches support multi-environment pipelines and can integrate automated testing into your release process. ### Option 1: Prophecy-native CI/CD Prophecy includes a native Git-based CI/CD workflow integrated directly into the project editor. This allows you to manage the entire lifecycle without leaving the Prophecy interface. In Prophecy's Git-based workflow, a release marks a specific version of your project by creating a Git tag, while deployment builds and pushes that version to your chosen environment. These steps usually run together but can also be executed independently. As part of the release process, Prophecy automatically builds the code, runs unit tests, and packages everything needed (such as JARs or wheels) to deploy pipelines and jobs. See Also * [Git](/data-engineering/ci-cd/git/git) to learn more about the Git workflow * [Unit tests](/data-engineering/ci-cd/tests) for validating pipeline functionality * [Deployment](/data-engineering/ci-cd/deployment/deployment) to deep dive into the phases of project deployment ### Option 2: External CI/CD with PBT If your organization already uses an external CI/CD system, you can integrate Prophecy projects using the Prophecy Build Tool (PBT), a command-line interface designed for automation. PBT works with systems like [GitHub Actions](/data-engineering/ci-cd/prophecy-build-tool/pbt-github-actions) and [Jenkins](/data-engineering/ci-cd/prophecy-build-tool/pbt-jenkins) to build and deploy Prophecy projects from a Git repository. This approach supports the same multi-environment model as native CI/CD. Use the `--fabric-ids` flag in your CI/CD configuration to target specific fabrics during deployment. To use PBT in your CI/CD pipeline: 1. Install the Prophecy Build Tool. 2. Configure secrets to securely store credentials and environment connection details. 3. Set up your GitHub Actions or Jenkins workflow to include build and deploy steps using PBT. For detailed instructions and examples, see [Prophecy Build Tool (PBT)](/data-engineering/ci-cd/prophecy-build-tool/prophecy-build-tool). # Unit tests for Data Engineering Source: https://docs.prophecy.ai/data-engineering/ci-cd/tests Implementing unit tests in Prophecy Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Writing good unit tests is one of the key stages of the CI/CD process. It ensures that the changes made by developers to projects will be verified and all the functionality will work correctly after deployment. Prophecy makes the process of writing unit cases easier by giving an interactive environment via which unit test cases can be configured across each component. There are two types of unit test cases which can be configured through Prophecy UI: 1. Output rows equality 2. Output predicates Let us understand both types in detail: ## Output rows equality Automatically takes a snapshot of the data for the component and allows to continuously test that the logic performs as intended. This would simply check the equality of the output rows. ### Example In the below example we would create below unit tests: 1. To check the join condition correctly for one-to-one mappings. 2. To check the join condition correctly for one-to-many mappings. ## Output predicates These are more advanced unit tests where multiple rules need to pass in order for the test as a whole to pass. Requires Spark expression to be used as predicates. ### Example In the below example we will create below unit tests: 1. Check that the value of amount column is `>0`. 2. Check whether first name is not equal to last name. ## Generating sample data for test cases automatically To generate sample input data automatically from the source DataFrame, this option can be enabled while creating unit test. Pipeline needs to run once, to generate units test based on auto-generated sample data. Let's generate sample data automatically for the unit test case we created in above example. ## Generated code Behind the scenes, the code for unit tests is automatically generated in our repository. Let's have a look at the generated code for our unit test above. ## Renaming the name of unit test # Data Engineering best practices Source: https://docs.prophecy.ai/data-engineering/development/best-practices Learn how to best use Prophecy for data engineering projects Applicable to the [Enterprise Edition](/data-engineering/administration/platform/editions) only. ## Projects Limit the total number of pipelines per project to keep your project modular. This helps you have: * Better Git version control and Git tagging * A faster code generation and compilation as common entities are compiled across all pipelines when changes are made * More control during deployment * Shared resources across teams and in [Package Hub](/data-engineering/extensibility/package-hub/package-hub) ## Re-usable entities Keep common entities in a common project. Common entities can include user defined functions, reusable-subgraphs, gems, and fully configurable pipelines. This allows you to share the project in a read-only and version controlled manner with other teams when you publish it as a [package](/data-engineering/extensibility/package-hub/package-hub). ## Pipelines 1. Limit the number of gems per pipeline to keep your pipeline modular. This helps you have: * Shorter recover and retry times for failed tasks in Spark pipelines * More control during orchestration * Shorter recovery times for failed jobs 2. Use [Job Sampling](/data-engineering/development/pipelines/pipeline-settings#job) only for debugging purposes and for smaller pipelines because sampling incurs a large computational penalty in Spark. ## Configurations You can configure Prophecy gems to use [configuration](/data-engineering/development/pipelines/configuration) variables such as path or a subset of path variables in Source and Target gems. Typically, configuration variables are static. However, if you want to assign a dynamic value to a configuration variable at runtime, you may overwrite the `Config` variable using a Script component. Run the following syntax at a low phase value (-1) before the rest of your pipeline: ```shell theme={null} Config.var_name =new_value ``` This lets you achieve a **dynamic runtime configuration.** ## Datasets Don't duplicate your dataset in the a pipeline. Your dataset contains a unique set of properties. If you duplicate it in your pipeline, its properties become unstable due to the duplicate copy of properties. ## Pipeline optimizations To optimize your pipeline: 1. For most cases, use the [Reformat gem](/data-engineering/gems/transform/reformat) instead of the [SchemaTransform gem](/data-engineering/gems/transform/schema-transform). The Reformat gem calls the Spark `select()` function once, which is not very computationally expensive. The SchemaTransform gem uses the Spark `withColumn()` function, which is applied to each column one-by-one (more computationally intensive). You should use the SchemaTransform gem if you are creating a small number of columns that will be specifically used for downstream calculations in subsequent gems. 2. If you have a gem that has multiple output ports, try caching the data in that gem. Spark lazily evaluates action calls, which means it reevaluates the same part of the flow unless you cache it. This is helpful before you branch to multiple output ports. Larger datasets may be too large to cache. 3. Broadcast smaller tables in your [Join gem](/data-engineering/gems/join-split/join) to increase your performance. Control the broadcast threshold based on your cluster size by setting the `spark.sql.autoBroadcastJoinThreshold` property to a value greater than 10MB. To learn more, see [Performance Tuning](https://spark.apache.org/docs/latest/sql-performance-tuning.html). 4. Remove the [OrderBy](/data-engineering/gems/transform/order-by) and [Deduplicate](/data-engineering/gems/transform/deduplicate) gems wherever you don't need them. If you need the Deduplicate gem, be mindful on which `Row to keep` to select.
The `first` and `last` options are more expensive than `any`. 5. Set an appropriate value for the `spark.sql.shuffle.partitions` property. For skewed, overparititioned, or underpartitioned Source datasets, consider using the [Repartition](/data-engineering/gems/join-split/repartition) gem to repartition your dataset to an appropriate number of partitions. # Spark Copilot Source: https://docs.prophecy.ai/data-engineering/development/copilot/copilot See how Copilot can help you build your pipeline Applicable to the [Enterprise Edition](/data-engineering/administration/platform/editions) only. Prophecy's Spark Copilot provides suggestions from an AI model as you develop your data pipelines and models. You can view and incorporate suggestions directly within the Prophecy visual editor and code editor. The Spark Copilot fetches context from environment metadata only. The Spark Copilot **does not** use knowledge graphs, which additionally store metadata relationships. ## Text to pipelines Get started on a new pipeline quickly by typing your prompt into the text box and Data Copilot will generate a new pipeline or modify an existing one. ### Start a new pipeline You can use Data Copilot to start a new pipeline by typing a simple English text prompt. Start a pipeline The following example uses Data Copilot to help start a pipeline: 1. Type a prompt with English text, such as `Which customers shipped the largest orders this year?` 2. If you'd like, review the suggested changes before you decide to keep or reject the suggested pipeline. Then interactively execute it to see the results. 3. View Data Copilot's suggested changes in the visual editor. ### Modify an existing pipeline You can also call Data Copilot to modify an existing model. Type a new text prompt, and Data Copilot will suggest a new sequence of data transformations. You don't necessarily have to select where you want to make your modification for Data Copilot to make its suggestion. Added/updated gems are highlighted in yellow. ## Next-transformation suggestions Data Copilot can suggest the next transformation in a series or the next expression within a gem. ### Suggest gems Data Copilot can suggest the next transformation for Leaf Nodes in a graph. Suggest gems See the following Join suggestion example: 1. Select and drop a dataset of interest on the canvas. 2. Data Copilot suggests datasets which are frequently used with the selected dataset. 3. Data Copilot then suggests a next transformation, in this case, a Join gem. ### Suggest Expressions As we continue development within gems, Data Copilot can suggest expressions within gems. Suggest expressions Within our [advanced Expression Builder](/data-engineering/gems/expression-builder) you can: 1. Type an English text prompt. 2. Data Copilot generates a code expression for a particular column. 3. Review the code expression, and if you'd like, try again with a different prompt. 4. Run the pipeline up to and including this gem, and observe the resulting data sample. ## Generate with AI Data Copilot can generate script gems, user-defined functions in Spark, or macro functions in SQL. ## Map with AI You don't have to worry about mapping the schema across your model. Data Copilot will map the target schema with the existing gems and datasets. ## Code with AI In addition to the visual editor above, you'll also see code completion suggestions in the code editor. Data Copilot helps you build your model in the code interface by making predictions as you type your code. And when you go back to the visual interface, you'll see your code represented as a model. ## Fix with AI If there are any errors in your gems, perhaps introduced upstream without your knowledge, Data Copilot will automatically suggest one-click fixes. The Fix with AI option appears on the diagnostic screen where you see the error messages or directly with the expression itself. ## Auto Documentation Understanding data assets is much easier with Data Copilot's auto-documentation. Data Copilot delivers summary documentation suggestions for all aatasets, pipelines, models, and orchestrations. ### Explain gems Here Data Copilot provides a high-level summary of a pipeline and more detailed description of each gem. ### Describe Datasets and Metadata How did a dataset change? Data Copilot recommends a description of the change for every edit you make. How was a column computed? Data Copilot suggests a plain English description that explains data sources and how every column is generated and what it represents. This is a big time saver! You can edit the documentation suggestions and commit them to your repository. ### Write Commit Messages and Release Notes Data Copilot auto-documents anywhere you need it - from the granular data sources and columns to gem labels, all the way to project descriptions. Copilot even helps you write commit messages and release notes. ## Data Tests and Quality Checks Unit tests and data quality checks are crucial for pipeline and job productionalization, yet many teams leave little time to develop these tests or worse, don't build them at all. With Data Copilot, you'll have one or more suggested [unit tests](/data-engineering/ci-cd/tests) that can be seamlessly integrated into your CI/CD process. Data Copilot also suggests data quality checks based on the data profile and expectations. # Data exploration for Data Engineers Source: https://docs.prophecy.ai/data-engineering/development/data-explorer/data-explorer Inspect interim data samples at each stage of your pipeline The Data Explorer helps you inspect interim data samples at each stage of your pipeline. By checking column structure, reviewing sample values, and confirming data types, you can catch issues early and ensure your pipeline is working as expected. ## Open the Data Explorer To use the Data Explorer, you need to [run](/data-engineering/development/runs/execution#interactive-execution) your pipeline to generate data samples. Click on any data sample in your pipeline to open the Data Explorer. Data sample in a pipeline ## Leverage the Data Explorer In the Data Explorer, you can: * Sort data by columns * Filter rows by specific values * Search across all values * Show or hide columns * Export the sample as CSV or JSON file * Save the transformation as a new gem Data explorer ## View complete dataset The Data Explorer loads a sample of your data by default. When you sort, filter, or search, these actions apply only to the visible rows in the sample. To work with the full dataset, do one of the following: * Click **Load More** at the bottom of the table until all rows are visible. * Click **Run** in the top-right corner of the preview. This refreshes the view and applies sorting and filtering to the entire dataset. # Data profiling for Data Engineers Source: https://docs.prophecy.ai/data-engineering/development/data-explorer/data-profile See high level statistics for data samples in your pipeline Data profiling allows you to view statistics on interim datasets in your pipeline. When you open a dataset's profile in the [Data Explorer](/data-engineering/development/data-explorer/data-explorer), you can visualize value distributions and data completeness to ensure your data matches expectations. ## Prerequisites To view data profiles, you need to: * Work on a PySpark project. * Upgrade the ProphecyLibsPython dependency 1.9.40 or later. * Use [selective data sampling mode](/data-engineering/development/runs/data-sampling#selective-sampling) in the pipeline. ## Quick profile The Data Explorer includes data profiles that are generated on your sample data. You'll be able to see high-level statistics for each column, including: * **Percent of non-blank values:** The percentage of values in the column that are not blank. * **Percent of null values:** The percentage of values in the column that are null. * **Percent of blank values:** The percentage of values in the column that are blank. * **Most common values:** Displays the top four most frequent values in the column, along with the percentage of occurrences for each. To view these statistics for your sample data, click **Profile** in the Data Explorer. Quick profile ## Expanded profile When you open the Data Explorer, you'll only see the data profile of the data **sample**. When you load the expanded data profile, Prophecy generates a more in-depth analysis on **all of the records** in the interim dataset. Expanded profile The full profile displays the following information: * **Data type**: The data type of the column. * **Unique values**: The number of unique values in the column. * **Longest value**: The longest value in the column and its length. * **Shortest value**: The shortest value in the column and its length. * **Most frequent value**: The most frequent value in the column and its number of occurrences. * **Least frequent value**: The least frequent value in the column and its number of occurrences. * **Minimum value**: The minimum value in the column. * **Maximum value**: The maximum value in the column. * **Average value length**: The average length of each value in the column. * **Null values**: The percent and number of null values in the column. * **Blank values**: The percent and number of blank values in the column. * **Non-blank values**: The percent and number of non-blank values in the column. * **Data summary**: An overview of the most common values in the column. You can click between columns in the expanded profile for quick access. ### Open expanded profile To view the expanded profile: 1. Click the dropdown arrow on the column you want to expand. 2. Select **Show Expanded Profile**. Show Expanded Profile # Datasets Source: https://docs.prophecy.ai/data-engineering/development/dataset Use datasets in your Spark project Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. In Prophecy, datasets are grouped by projects and rely on the following: * **Schema**: The structure or shape of the data, including column names, data types, and the method for reading and writing the data in this format. * **Fabric**: The execution environment in which the data resides. ## Create datasets Datasets are created where they are first used in a [Source or Target gems](/data-engineering/gems/gems). A dataset definition includes its: * **Type**: The type of data you are reading/writing like CSV, Parquet files or catalog tables. * **Location**: The location of your data. It could be a file path for CSV or a table name. * **Properties**: Properties consists of Schema and some other attributes specific to the file format. For example, in case of CSV, you can give Column delimiter in additional attributes. You can also define Metadata for each column here like description, tags, and mappings. Datasets can be used by any pipeline within the same project, and in some cases by other projects within the same team. ## View datasets There are two ways to view a list of datasets: * To see all datasets, navigate to **Metadata > Datasets**. * To see only one project's datasets, navigate to **Metadata > Projects**. Then, open a project. Click on the **Content** tab, and then the **Datasets** subtab. ## Dataset Metadata If you open the metadata page for one of the datasets, you'll find the following information: | Name | Description | | ------------------- | ------------------------------------------------------------------- | | Dataset name | The name of this dataset, which is editable. | | Dataset description | The description of this dataset, which is editable. | | Dataset properties | A subset of properties used for reading or writing to this Dataset. | | Dataset schema | The columns of this dataset and their data types. | | Delete Dataset | The option to delete this dataset. Use with caution. | In the **Relations** tab, there is additional information about where and how this dataset is used. | Name | Description | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | | Physical Datasets | Location of the dataset in relation to a fabric. | | Pipelines | A list of pipelines that use this dataset, with the `Relation` column indicating if it is for `Read` or `Write` purposes. | | Jobs | A list of jobs that use this dataset, with the `Relation` column indicating if it is for `Read` or `Write` purposes. | | Open Lineage Viewer | The option to open this dataset in the [Lineage](/data-engineering/lineage/lineage) viewer, showing column-level lineage for this dataset. | ## Publishing and sharing datasets As part of the project release process, datasets within that project are *published* to other projects within the same Team, and can be published to other Teams in read-only mode. This allows you to share your dataset configurations with other Teams without allowing them to make changes to the original dataset definitions. Let's see this in action: 1. `DI_TEAM` is the central Data Infrastructure team. They have defined a common project named `DI_Common_Python`. 2. `DI_Common_Python` has a number of datasets defined within it: DI Common Datasets 3. The `DI_Team` merges and releases the `DI_Common_Python` project, tagging it `0.1`. DI Common Release 4. As you can see, the `DI_Team` has published the `DI_Common_Python` project to the `DE_Team`, the Data Engineering Team. 5. Now, whenever the `DE_Team` builds pipelines, they can see the following: Common Datasets We can see the `DI_Common_Python` project's datasets, and the fact that they're listed as `Read-only`. This means that `DE_Team` can *use* the datasets, but cannot *edit* them. For regular usage, we suggest having only one instance of a particular dataset within a pipeline, as the dataset's properties and underlying data can change each time the dataset is read or written. # Business rules Source: https://docs.prophecy.ai/data-engineering/development/functions/business-rules-engine/business-rules-engine Use business rules to automate business decisions Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Business rules empower organizations to model, manage, and automate repeatable business decisions throughout the enterprise. ## Overview The business rules engine in Prophecy lets you incorporate business logic in your Pipelines. Often, different users will interact with business rules at various stages of Pipeline development. One common workflow is as follows: 1. Business users create business rules using predefined enterprise logic. 2. These business rules are deployed and become available as [dependencies](/data-engineering/extensibility/dependencies/spark-dependencies) in the [Package Hub](/data-engineering/extensibility/package-hub/package-hub). 3. Data engineers and others working on Pipelines can incorporate these rules using the [SchemaTransform gem](/data-engineering/gems/transform/schema-transform). These stages are explained in more detail below. ## Business rule parameters Business rules require the following parameters. | Field | Description | | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | | Input Columns | The columns you would like to use in your business rule conditions. Each column must have a Name and Type defined. | | Output Columns | The names of the output columns that the business rule will calculate. | | Rules | The set of conditions that define the business rule. These can be grouped by date ranges, such that they only apply during a certain time period. | ## Create business rules To create a new business rule: 1. In the project browser, click the **+** icon next to **Functions**. 2. Name the Rule and choose **Business Rule** as the **Function Type**. Then, click **OK**. Create business rule Next, you need to define the rule parameters. 1. Define columns in the **Input Columns** table using a name, type, and (optionally) description. 2. Define columns in the **Output Columns** table with an optional description. 3. Each column in these tables will automatically be added to the **Rules** table. Rule Parameters Each row in the Rules table corresponds to one rule condition. To add a rule condition: 1. Add conditions under the corresponding inputs in SQL expression format. Note that each expression must be true in a row (logical `AND`) for the rule condition to be met. 2. Add an output value for your condition. This can be a hard-coded value, or you can write in SQL expression format. The default output value is `Null`. 3. Optionally, you can add a description of the rule. Prophecy will generate an error if there is conflicting logic in your conditions. ### Example: IdentifyHighSpendingCustomer Take a look at the business rule in the following image. IdentifyHighSpendingCustomer Rule In this example, we created an **IdentifyHighSpendingCustomer** business rule. This rule represents the following statement: > If a customer's spending is more than $500, or their total order amount is more than $500, then they are a high spending customer. Otherwise, they are not a high spending customer. As you can see, we needed to use multiple rule conditions to achieve this outcome. Additionally, you can see that the output is either `1` or `0`. This is because we decided to represent whether a customer was a high spender or not with a binary flag. ## Share business rules You can also import business rules into projects via Packages. Imported rules are read-only and can only be edited from their source project. This can be useful if: * You want to group rules by their function or use case. * You only want specific users to create and edit rules. * You want to reuse rules in multiple projects. ### Example: PromoCodeRule Let's say you want to create a PromoCodeRule that will be used in various other projects. 1. Start by creating a project where you will define the business rule. 2. Add the business rule to the project. 3. Commit your changes to the project. 4. Merge the changes to the main branch. 5. [Release and deploy](/data-engineering/ci-cd/deployment/deployment) the project. Then, you must give other users access to your project. 1. In your project metadata, open the **Access** tab. 2. Toggle-on the option to **Publish to Package Hub**. This will make the Package available to others. When someone adds the Package as a [dependency](/data-engineering/extensibility/dependencies/spark-dependencies) in their project, they will be able to see the rule definition. However, they will not be able to edit the fields. PromoCodeRule This example rule includes a set of conditions to determine the type of promotions that a customer is eligible for. ## Use business rules in your Pipeline To use a business rule in your Pipeline, you can use the [SchemaTransform gem](/data-engineering/gems/transform/schema-transform). 1. Add a SchemaTransform gem to the Pipeline. 2. Open the gem and add the appropriate input. 3. Click **Add Transformation**. 4. In the **Operation** dropdown, choose **Add Rule**. 5. Choose the appropriate rule in the **Rule** field. This will populate the **New Column** field. If an input column has the same name as the new column, then its data will be overwrittenβ€”no new column will be appended. After adding a business rule, Prophecy will automatically perform a few checks to verify that: * Each rule input column exists in the gem input. * The type of each rule input column matches that of the gem input. Error messages can be found in the Diagnostics of the gem. You can add multiple business rules to the SchemaTransform gem at a time. You can also use the output column of one rule as an input column for a subsequent rule. ## View the business rules in code Prophecy automatically compiles visually-developed business rules into code. Business rules are stored in the **functions** folder of your pipeline's code. This is true for both Python and Scala projects. Note that you can also see the imported business rules in the code view. Business rules in Python and Scala # User-defined functions Source: https://docs.prophecy.ai/data-engineering/development/functions/user-defined-functions Create and import UDFs in your pipeline Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Prophecy lets you create and import user-defined functions (UDFs), which can be used anywhere in the pipeline. Prophecy supports creating UDFs written in Python/Scala and importing UDFs written in SQL. | Project Type | Create UDFs | Import UDFs | | :----------- | :----------- | :------------ | | Python | Python/Scala | SQL | | Scala | Python/Scala | Not supported | Learn about UDF support in Databricks on our documentation on cluster [access modes](/data-engineering/fabrics/spark-provider/databricks/UCShared). ## Create UDFs Prophecy supports creating UDFs written in Python or Scala. ### Parameters | Parameter | Description | Required | | :---------------------- | :------------------------------------------------------------------------------------------------------------------------------------------ | :------- | | Function name | The name of the function as it appears in your project. | True | | UDF Name | The name of the UDF that will register it. All calls to the UDF will use this name. | True | | Definition | Definition of the UDF function.
For example, `udf((value:Int)=>value*value)` | True | | UDF initialization code | Code block that contains initialization of entities used by UDFs. This could, for example, contain any static mapping that a UDF might use. | False | ### How to Create UDFs 1. Create a new function. You can find the **Functions** section in the left sidebar of a project page. Add a function to the pipeline 2. Define the function. Define the function 3. Call the function. Call the function ```python theme={null} country_code_map = {"Mexico" : "MX", "USA" : "US", "India" : "IN"} def registerUDFs(spark: SparkSession): spark.udf.register("get_country_code", get_country_code) @udf(returnType = StringType()) def get_country_code(country: str): return country_code_map.get(country, "Not Found") ``` ```scala theme={null} object UDFs extends Serializable { val country_code_map = Map("Mexico" -> "MX", "USA" -> "US", "India" -> "IN") def registerUDFs(spark: SparkSession) = spark.udf.register("get_country_code", get_country_code) def get_country_code = udf { (country: String) => country_code_map.getOrElse(country, "Not Found") } } ``` ## Import UDFs UDFs from the Unity Catalog are automatically available in Python projects when you attach to the relevant Databricks fabric. You can call these UDFs from any gem in the project. Call SQL function To view a function, open the Environment browser and expand the Catalog and Schema containing the function. Refresh the project editor to access any new functions added in Databricks. View SQL function ## UDFs across pipelines User-defined functions (UDFs) are defined at the project level, so they are shared across all pipelines in the project. However, each pipeline keeps its own local copy of the UDF code. Prophecy updates this copy only when you open the pipeline. So if someone edits or adds a UDF in one pipeline, those changes won't automatically appear in other pipelines until you open them. At that point, Prophecy copies the latest UDF definitions into the pipeline, and you'll see them as uncommitted changes in the code view. # Model configurations Source: https://docs.prophecy.ai/data-engineering/development/models/configuration Configure SQL project and model variables Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Model configurations are settings that define how a model should be built and behave within your data warehouse. When you open a SQL project, you can find **Configuration** in the project settings menu beside the project name. If you use a configuration in your model, you can switch to the code view to see the configuration encoded in the `dbt_project.yml` or `schema.yml/properties.yml` file. See also Further information can be found in the dbt documentation on [model configurations](https://docs.getdbt.com/reference/model-configs). ## Types Configurations are variables that you can use in various gem fields. There are two types of configurations. * **Model configurations**: Only accessible in a specific model. * **Project configurations**: Accessible to any component within a specific project. ## Syntax The variable name and value should both be valid in Python. The way you reference these variables differ between model and project configurations. The table below shows some usage examples for each type of configuration. | Type | Python Syntax | SQL Syntax | | ------- | ---------------------------- | ---------------------------------- | | Model | `key` | `{{ key }}` | | Project | `var("key", "defaultvalue")` | `{{ var("key", "defaultvalue") }}` | Note that the `defaultvalue` is optional for project configurations. # What are SQL models? Source: https://docs.prophecy.ai/data-engineering/development/models/models Models define a single target table or view in a SQL warehouse Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. In Prophecy, a model comprises a set of gems that process data into one output. In other words, each model corresponds to a **single table** in your database. Models leverage the dbt build system and can run on either SQL fabrics or Prophecy fabrics. To work with the dbt build, Prophecy saves each visual model as a SQL file in your project repository in Git. Prophecy's visual interface supports SQL models only; if you'd like to define Python models, you must do so using the code interface. ## Create models To add a new model to your project: 1. Open your project in the project editor. 2. Click **+ Add Entity** from the bottom of the **Project** tab in the left sidebar. 3. Click **Model**. 4. In the Add Model dialog, add a **Model Name**. 5. Review the path where the model will be saved in the project repository. In most cases, the default `model` path is sufficient. 6. Click **Create**. This opens a new model canvas that is prepopulated with a target model. Note that the dbt framework restricts models to one target output. While you can develop models visually using gems, you can also write models directly in the code view, which is automatically synced with the visual view. ## Compatible gems Each gem in a model maps to a SQL statement. As you configure gems on the visual canvas, Prophecy automatically generates the corresponding SQL, determines whether to use a CTE or subquery for each step, and integrates your changes into the overall model. ## Advanced settings The **Advanced Settings** dialog lets you set dbt configurations at the project, folder, or model level. Each setting corresponds to standard dbt configurations, typically defined in a dbt project [YAML file](https://docs.getdbt.com/docs/build/projects#project-configuration). They cover properties such as materialization behavior, physical storage, metadata, and access controls. Some settings reflect properties found in the target model gem. When you update a setting through the Advanced Settings panel, it automatically syncs with the corresponding target model gem when relevant. To open the Advanced Settings: 1. Open the project settings menu beside the project name. 2. Select **Advanced Settings**. 3. Choose to edit the Project Settings, Folder Settings, or Model Settings. Advanced Settings Folder Settings only apply to directories that contain models. ## Schedule models To schedule automated model execution: * The project must use the [Normal Git Storage Model](/data-analysis/development/versioning/version-control). * You need to use an external orchestrator, such as [Databricks Jobs](/data-engineering/orchestration/databricks-jobs) or Apache Airflow DAGs. ## Models vs pipelines Models and pipelines are two different SQL project components. The following table describes the key differences between models and pipelines. | Feature | Models | Pipelines | | ---------------- | ---------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | | Execution Engine | Models run entirely on the SQL Warehouse. | Pipelines run on the SQL Warehouse and can also use Prophecy Automate. | | Supported Gems | Models support only gems that execute within the SQL Warehouse. | Pipelines support additional gems that run in Prophecy Automate, such as the Email or REST API gem. | | Data Sources | Models can only use native tables/models as sources and targets. | Pipelines can use both native tables/models and external data sources. | | Outputs | Models are limited to a single output. | Pipelines can write multiple outputs. | | Orchestration | Models must be orchestrated using Databricks Jobs or Airflow. | Pipelines can be orchestrated externally or natively using Prophecy Automate. | | Exportability | Models generate SQL code that can be run outside of Prophecy. | Pipelines that include Prophecy Automate gems cannot be run outside of Prophecy. | ### Show underlying models Many visual transformations in pipelines are compiled into models under the hood. If you are working on a pipeline, you can view and edit the code of underlying dbt models in a pipeline. However, you cannot visually edit these underlying models. To view these models, select **Show Models** from the project interface. Show Models # Dynamic target location Source: https://docs.prophecy.ai/data-engineering/development/models/sources-target/location Use a dynamic location for your target model Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. By default, Prophecy writes your target model to the database and schema defined in the attached fabric. You can update the location of the target model in the **Location** tab of the gem dialog. This page includes an example of how you can make the write location of the table dynamic. ## Use case Assume you want to use one database location during development and interactive execution, but you want to write to a different database for schedules running in production. You can use a configuration variable to do so. ### Create the variable First, you will have to create the variable: 1. Open the project settings menu beside the project name and select **Configuration**. 2. Make sure you are in **Project Configuration**. 3. Create a variable and make the default value the name of the development database you want to use during interactive execution. The value should be a string. ### Overwrite the default model location Next, you need to add the variable to your target model location. 1. Open a target model gem. 2. Click on the **Location** tab. 3. Enable the **Overwrite** toggle for the database. 4. Click on **Advanced Mode** on the right side of the database field. 5. From the dropdown that appears, select **Configuration Variable**. 6. Choose the configuration variable you created in the previous section. 7. **Save** your changes. Location ### Assign the variable a value Then, let's change the variable to save to a **production** database. 1. Create a job that includes your model. 2. Open the model configuration and add the **Supply variables to project** dbt property. 3. Add your project variable and assign it the name of the production database. This will override the default value provided when you configured the variable. Now, when the job runs, your model should be stored in the production database. # Model sources and targets Source: https://docs.prophecy.ai/data-engineering/development/models/sources-target/sources-target Use models to read and write data Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Model sources and targets vary slightly from those of a pipeline. The primary difference is that all model sources and targets must point to tables in the SQL warehouse. ## Sources When you create a new model, you need to define an input data source. The data source can be: * Another model. You can drag a model from the Project tab of the left sidebar onto your canvas to use it as a source. * A [Table gem](/data-analysis/gems/source-target/source-target). You can either use preconfigured tables from the Project tab of the left sidebar, or you can browse SQL warehouse tables in the Environment tab of the left sidebar. ## Targets Target models let you define how you want to materialize your data using write formats. When you open a target model configuration, you'll see the following tabs: * **Type & Format**: Update the format of the model between different table materialization types. * **Location**: Update the location by overwriting the database, schema, or alias. * **Schema**: Make schema changes and set optional dbt properties. * **SQL Query**: Enable and create a custom SQL query to include at the end of the target model. * **Write Options**: Choose a specific write mode such as overwrite, append, and merge. Target Model tabs # BigQuery target models Source: https://docs.prophecy.ai/data-engineering/development/models/target-platforms/bigquery-target Configure target models for BigQuery SQL Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. To configure a target model that will be written to BigQuery, reference the following sections. ## Type & Format BigQuery supports the following materialization types for target models. The type determines the underlying physical format of your target model. | Materialization type | Description | | -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | View (default) | Rebuilt as a view on each run. Views always reflect the latest source data but don't store any data themselves. | | Table | Rebuilt as a table on each run. Tables are fast to query but can take longer to build. Target tables support multiple write modes. | | Ephemeral | Not built in the database. The model's logic is inlined into downstream models using a common table expression (CTE). Use for lightweight transformations early in your DAG. | | Materialized View | Acts like a hybrid of a view and a table. Supports use cases similar to incremental models. Creates a materialized view in the target warehouse. | ## Location Review the location where your model will be written. Any changes you make to the **Overwrite location** section will be reflected in the **Location** that Prophecy generates. | Location Parameter | Description | Advanced mode | | --------------------------- | --------------------------------------------------------------------------------------------- | ------------- | | Project ID | Google Cloud project where the model will be built. | Yes | | Database (BigQuery dataset) | Name of the BigQuery dataset where the model will be created. Acts as the schema in BigQuery. | Yes | | Alias | Sets the name of the resulting table or view. Defaults to the model name if not specified. | No | ## Schema Define the schema of the dataset and optionally configure additional properties. The schema includes column names, column data types, and optional column metadata. When you expand a row in the Schema table, you can add a column description, apply column tags, and enable/disable quoting for column names. ### Properties Each property maps to a certain dbt configuration that may be generic to dbt or specific to a platform like BigQuery. If you do not add a property explicitly in the Schema tab, Prophecy uses the dbt default for that property. | Property | Description | Config type | | --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------- | | Dataset Tags | Add tags to the dataset. These tags can be used as part of the resource selection syntax in dbt. | Generic | | Contract Enforced | Enforce a [contract](https://docs.getdbt.com/docs/mesh/govern/model-contracts) for the model schema, preventing unintended changes. | Generic | | Show Docs | Control whether or not nodes are shown in the [auto-generated documentation website](https://docs.getdbt.com/docs/build/view-documentation). | Generic | | Enabled | Control whether the model is included in builds. When a resource is disabled, dbt will not consider it as part of your project. | Generic | | Meta | Set metadata for the table using key-value pairs. | Generic | | Group | Assign a group to the table. | Generic | | Persist Docs Columns | Save column descriptions in the database. | Generic | | Persist Docs Relations | Save model descriptions in the database. | Generic | | Cluster By | [Cluster data in the table](https://cloud.google.com/bigquery/docs/clustered-tables) by the values of specified columns to improve query performance and reduce costs. | BigQuery | | Partition Expiration Days | If using date or timestamp partitions, this property defines the number of days from the partition date to expiration. | BigQuery | | Require Partition Filter | Requires anyone querying this model to specify a partition filter, otherwise their query will fail. | BigQuery | | Time Ingestion Partitioning | Enables partitioning based on when data is ingested into the table, using [BigQuery's](https://cloud.google.com/bigquery/docs/partitioned-tables#ingestion_time) `_PARTITIONTIME` column. | BigQuery | For more detailed information, see the [dbt reference documentation](https://docs.getdbt.com/reference/references-overview). ## SQL Query Add a custom SQL query at the end of your target model using the BigQuery SQL dialect. This allows you to apply a final transformation step, which can be useful if you're importing an existing codebase and need to add conditions or filters to the final output. Custom queries support Jinja, dbt templating, and [variable](/data-engineering/development/models/configuration) usage for your last-mile data processing. You can reference any column present in the list of input ports beside the SQL query. You can only add additional input portsβ€”the output port cannot be edited. ## Write Options For a complete guide to defining how to write target tables, visit [Write strategies](/data-analysis/gems/source-target/table/write/write-options). ## Data Tests A data test is an assertion you define about a dataset in your project. Data tests are run on target models to ensure the quality and integrity of the final data that gets written to the warehouse. Learn how to build tests in [Table tests](/data-analysis/development/tests/table-tests). # Databricks target models Source: https://docs.prophecy.ai/data-engineering/development/models/target-platforms/databricks-target Configure target models for Databricks SQL Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. To configure a target model that will be written to Databricks, reference the following sections. ## Type & Format Databricks supports the following materialization types for target models. The type determines the underlying physical format of your target model. | Materialization type | Description | | -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | View (default) | Rebuilt as a view on each run. Views always reflect the latest source data but don't store any data themselves. | | Table | Rebuilt as a table on each run. Tables are fast to query but can take longer to build. Target tables support multiple write modes. | | Ephemeral | Not built in the database. The model's logic is inlined into downstream models using a common table expression (CTE). Use for lightweight transformations early in your DAG. | | Materialized View | Acts like a hybrid of a view and a table. Supports use cases similar to incremental models. Creates a materialized view in the target warehouse. | ## Location Review the location where your model will be written. Any changes you make to the **Overwrite location** section will be reflected in the **Location** that Prophecy generates. | Location Parameter | Description | Advanced mode | | ------------------ | ------------------------------------------------------------------------------------------ | ------------- | | Catalog | Catalog where the model will be created. | Yes | | Schema | Schema inside the catalog where the model will be created. | Yes | | Alias | Sets the name of the resulting table or view. Defaults to the model name if not specified. | No | ## Schema Define the schema of the dataset and optionally configure additional properties. The schema includes column names, column data types, and optional column metadata. When you expand a row in the Schema table, you can add a column description, apply column tags, and enable/disable quoting for column names. ### Properties Each property maps to a certain dbt configuration that may be generic to dbt or specific to a platform like Databricks. If you do not add a property explicitly in the Schema tab, Prophecy uses the dbt default for that property. | Property | Description | Config type | | ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- | ----------- | | Dataset Tags | Add tags to the dataset. These tags can be used as part of the resource selection syntax in dbt. | Generic | | Contract Enforced | Enforce a [contract](https://docs.getdbt.com/docs/mesh/govern/model-contracts) for the model schema, preventing unintended changes. | Generic | | Show Docs | Control whether or not nodes are shown in the [auto-generated documentation website](https://docs.getdbt.com/docs/build/view-documentation). | Generic | | Enabled | Control whether the model is included in builds. When a resource is disabled, dbt will not consider it as part of your project. | Generic | | Meta | Set metadata for the table using key-value pairs. | Generic | | Group | Assign a group to the table. | Generic | | Persist Docs Columns | Save column descriptions in the database. | Generic | | Persist Docs Relations | Save model descriptions in the database. | Generic | | Clustered By | Each partition in the created table will be split into a fixed number of buckets by the specified columns. | Databricks | | Buckets | The number of buckets to create while clustering. Required if **Clustered By** is specified. | Databricks | For more detailed information, see the [dbt reference documentation](https://docs.getdbt.com/reference/references-overview). ## SQL Query Add a custom SQL query at the end of your target model using the Databricks SQL dialect. This allows you to apply a final transformation step, which can be useful if you're importing an existing codebase and need to add conditions or filters to the final output. Custom queries support Jinja, dbt templating, and [variable](/data-engineering/development/models/configuration) usage for your last-mile data processing. You can reference any column present in the list of input ports beside the SQL query. You can only add additional input portsβ€”the output port cannot be edited. ## Write Options For a complete guide to defining how to write target tables, visit [Write strategies](/data-analysis/gems/source-target/table/write/write-options). ## Data Tests A data test is an assertion you define about a dataset in your project. Data tests are run on target models to ensure the quality and integrity of the final data that gets written to the warehouse. Learn how to build tests in [Table tests](/data-analysis/development/tests/table-tests). # Snowflake target models Source: https://docs.prophecy.ai/data-engineering/development/models/target-platforms/snowflake-target Configure target models for Snowflake SQL Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. To configure a target model that will be written to Snowflake, reference the following sections. ## Type & Format Snowflake supports the following materialization types for target models. The type determines the underlying physical format of your target model. | Materialization type | Description | | -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | View (default) | Rebuilt as a view on each run. Views always reflect the latest source data but don't store any data themselves. | | Table | Rebuilt as a table on each run. Tables are fast to query but can take longer to build. Target tables support multiple write modes. | | Ephemeral | Not built in the database. The model's logic is inlined into downstream models using a common table expression (CTE). Use for lightweight transformations early in your DAG. | ## Location Review the location where your model will be written. Any changes you make to the **Overwrite location** section will be reflected in the **Location** that Prophecy generates. | Location Parameter | Description | Advanced mode | | ------------------ | ------------------------------------------------------------------------------------------ | ------------- | | Catalog | Catalog where the model will be created. | Yes | | Schema | Schema inside the catalog where the model will be created. | Yes | | Alias | Sets the name of the resulting table or view. Defaults to the model name if not specified. | No | ## Schema Define the schema of the dataset and optionally configure additional properties. The schema includes column names, column data types, and optional column metadata. When you expand a row in the Schema table, you can add a column description, apply column tags, and enable/disable quoting for column names. ### Properties Each property maps to a certain dbt configuration that may be generic to dbt or specific to a platform like Snowflake. If you do not add a property explicitly in the Schema tab, Prophecy uses the dbt default for that property. | Property | Description | Config type | | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------- | | Dataset Tags | Add tags to the dataset. These tags can be used as part of the resource selection syntax in dbt. | Generic | | Contract Enforced | Enforce a [contract](https://docs.getdbt.com/docs/mesh/govern/model-contracts) for the model schema, preventing unintended changes. | Generic | | Show Docs | Control whether or not nodes are shown in the [auto-generated documentation website](https://docs.getdbt.com/docs/build/view-documentation). | Generic | | Enabled | Control whether the model is included in builds. When a resource is disabled, dbt will not consider it as part of your project. | Generic | | Meta | Set metadata for the table using key-value pairs. | Generic | | Group | Assign a group to the table. | Generic | | Persist Docs Columns | Save column descriptions in the database. | Generic | | Persist Docs Relations | Save model descriptions in the database. | Generic | | Cluster By | Order and cluster the table by specified keys to reduce Snowflake's automatic clustering work. Accepts a string or list of strings. | Snowflake | | Transient | Write the table as a Snowflake [transient table](https://docs.snowflake.com/en/user-guide/tables-temp-transient#transient-tables), which does not preserve history and can reduce costs. By default, all Snowflake tables created by dbt are transient. | Snowflake | | Automatic Clustering | Enables automatic clustering if manual clustering is enabled for your Snowflake account. You do not need to use this property if automatic clustering is enabled by default on your Snowflake account. | Snowflake | | Snowflake Warehouse | Override the Snowflake warehouse that is used for this target model. | Snowflake | | Query Tag | Add [query tags](https://docs.snowflake.com/en/sql-reference/parameters#query-tag) to your model. | Snowflake | | Copy Grants | Enable to preserve any existing access control grants (like GRANT SELECT TO ROLE analyst) when rebuilding the table or view. | Snowflake | | Secure | Create a [secure view](https://docs.snowflake.com/en/user-guide/views-secure) that hides underlying data from unauthorized users. | Snowflake | | Target Lag | Defines how frequently the table should be [automatically refreshed](https://docs.snowflake.com/en/user-guide/dynamic-tables-target-lag). | Snowflake | For more detailed information, see the [dbt reference documentation](https://docs.getdbt.com/reference/references-overview). ## SQL Query Add a custom SQL query at the end of your target model using the Snowflake SQL dialect. This allows you to apply a final transformation step, which can be useful if you're importing an existing codebase and need to add conditions or filters to the final output. Custom queries support Jinja, dbt templating, and [variable](/data-engineering/development/models/configuration) usage for your last-mile data processing. You can reference any column present in the list of input ports beside the SQL query. You can only add additional input portsβ€”the output port cannot be edited. ## Write Options For a complete guide to defining how to write target tables, visit [Write strategies](/data-analysis/gems/source-target/table/write/write-options). ## Data Tests A data test is an assertion you define about a dataset in your project. Data tests are run on target models to ensure the quality and integrity of the final data that gets written to the warehouse. Learn how to build tests in [Table tests](/data-analysis/development/tests/table-tests). # Configurations Source: https://docs.prophecy.ai/data-engineering/development/pipelines/configuration Control how a pipeline behaves during execution Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. A configuration is a set of predefined variables and values that control how a data pipeline behaves during execution. By using configurations, you can dynamically adapt a pipeline to different environments or scenarios without modifying the pipeline itself. ## Configuration hierarchy Configurations are hierarchical and can be defined at the following levels: * **Project**: Shared across all pipelines in a project. * **Pipeline**: Overrides or extends project-level variables for a specific pipeline. * **Subgraph**: Local to a subgraph, with options to inherit from parent pipelines. * **Job**: Specifies the configuration to use when running pipelines in production. ## Project and pipeline configurations Each pipeline has access to two types of configurations: * **Project configuration**: Defines global variables that are shared across all pipelines in the project. * **Pipeline configuration**: Defines variables specific to the current pipeline and can override project-level values. To view or edit these configurations, click the **Config** button in the pipeline header. Each configuration includes two tabs: * [Schema](#schema-tab): Define the structure and data types for your configuration variables. * [Config](#config-tab): Set the values for each variable and create configuration instances for different environments. You can access the project configuration from any pipeline. Pipeline schema ### Schema tab In the Schema tab, you define the variables you want to use. The following table lists the information to define each variable. | Parameter | Description | | ----------- | ------------------------------------------------------------------------------------------------------------- | | Name | The unique identifier for the variable. | | Type | The data type assigned to the variable. Supported types are listed below. | | Optional | Indicates whether the variable is optional. If not selected, you must provide a default value for the config. | | Description | An optional text field to provide additional context or details about the variable. | Prophecy supports the following data types for configs. | Data type | Description | | ------------------ | --------------------------------------------------------------------------------------------------------------- | | `string` | A plain text value, entered via a single-line text input. | | `boolean` | A `true` or `false` value, selected from a dropdown. | | `date` | A calendar date in `dd-mm-yyyy` format, chosen using a date picker. | | `timestamp` | A specific date and time in `dd-mm-yyyyTHH:MM:SSZ+z` format (with time zone), selected using a datetime picker. | | `double` | A 64-bit floating-point number entered in a numeric field. | | `float` | A 32-bit floating-point number entered in a numeric field. | | `int` | A 32-bit integer entered in a numeric field. | | `long` | A 64-bit integer entered in a numeric field. | | `short` | A 16-bit integer entered in a numeric field. | | `array` | A list of values of the same type, added one by one in a multi-value input field. | | `record` | A structured object with multiple named fields, configured through a nested group of inputs. | | `secret` | A sensitive string (like a password or token), selected from your fabric's list of preconfigured secrets. | | `spark_expression` | A Spark SQL expression, written in a code editor with syntax highlighting. | ### Config tab The Config tab allows you to set values for the variables defined in your schema and create multiple configuration instances for different environments or use cases. Each configuration instance maintains its own set of values while inheriting the same schema structure. #### Configuration instances At the top of the Config tab, use the dropdown menu to switch between configuration instances. Every project and pipeline includes a `default` configuration that provides default values for all defined variables. You can create additional instances (such as `prod` for production) with alternate values, while reusing the same schema. To create a new configuration instance: 1. Open the config dropdown. 2. Click **New Configuration**. 3. In the **Instance Name** field, name your instance. 4. Click **Create**. 5. Add new values or maintain the default values from the `default` config. 6. Click **Save** to save the new values in the config. #### Set configuration values In the Config tab, you'll also see a form with all the variables you defined in the Schema tab. Each variable appears as an input field that matches its data type, for example: * Text fields for string variables * Numeric inputs for integer, float, and other numeric types * Dropdowns for boolean values * Date/time pickers for temporal data types If a variable in the schema is not optional, you must add a default value in the Config tab. #### Override project-level configs Pipelines inherit configuration values from the project-level `default` configuration. To change these values for a specific pipeline, create a new configuration instance for that pipeline. Within this instance, you can override any inherited values as needed. The `default` pipeline configuration always reflects the project-level default values. ## Use configuration variables When you want to call configuration variables in your pipeline, reference them using Jinja syntax. You can use the following syntax examples for accessing elements of array and record fields: * For a string: `{{ config_name }}` * For an array: `{{ config1.array_config[23] }}` * For a record: `{{ record1.record2.field1 }}` Jinja is enabled by default in new pipelines. You can disable Jinja support in **Pipeline Settings > Enable Jinja-based configuration**. ### Syntax for Source and Target gems To use configuration variables for specifying file locations for Source and Target gems, use the following syntax: `${config_variable_name}`. Jinja syntax is not supported for Source and Target gems. ### Syntax for different languages Depending on the Visual Language configured in your [Pipeline Settings](/data-engineering/development/pipelines/pipeline-settings), you can also use that language's syntax to call variables. | Visual Language | Syntax | Expression usage | | --------------- | -------------------- | -------------------------- | | SQL | `'$config_name'` | `expr('$config_name')` | | Scala | `Config.config_name` | `expr(Config.config_name)` | | Python | `Config.config_name` | `expr(Config.config_name)` | ## Runtime configuration After defining your configuration schema and setting up configuration instances, you need to specify which configuration to use when your pipeline actually runs. This process differs depending on the type of pipeline execution. ### Interactive execution For development and testing purposes, you can run pipelines interactively within the Prophecy interface. To control which configuration is used during these interactive runs: 1. Open the **Pipeline Settings** from the project editor. 2. Navigate to the **Run Settings** section. 3. Select the appropriate configuration from the list of existing configs. 4. **Save** your changes. Interactive runs use the `default` config by default. Choose config for interactive run ### Job execution When deploying pipelines to production environments, you typically run them as scheduled jobs rather than interactive executions. To specify your configuration values for jobs: 1. Add a Pipeline gem to the job canvas. 2. Open the Pipeline gem. 3. For the **Pipeline to schedule** field, choose the appropriate pipeline. 4. In the **Schema** tab, review the schema that has been inherited from the project and pipeline configs. 5. Change the selected project and pipeline configs, or keep the default. 6. If required, switch to the **Config** tab to override inherited values. Choose config for job execution ## Subgraph configurations [Subgraphs](/data-engineering/gems/subgraph/subgraph) can have their own dedicated configurations that control behavior within the subgraph's scope. This allows you to create reusable pipeline components with their own configurable parameters. To add configs to a subgraph: 1. Open the subgraph. 2. Click **+ Config**. 3. In the **Schema** tab, review the config schema. This might already included inherited project configuration variables. 4. To overwrite pipeline configuration variable definitions, click **Copy Pipeline Configs** in this tab. 5. Open the **Config** tab. 6. Define and save a set of values as a config. 7. Change the selected project and pipeline configs, or keep the default. 8. **Save** your changes. Subgraph configuration Upon creation, subgraph configurations will also be included in the pipeline configurations. Both the schema and config values for subgraphs can be edited from the pipeline configuration dialog. ## Code All configuration instances and values are automatically converted to code. 1. Open `Config.scala` in the `/config` folder. 2. View the default configuration code. 3. Find additional configurations that are packaged as JSON files in the `resources/config` folder. Config scala code 1. Open `Config.py` in the `/config` folder. 2. View the default configuration code. 3. Find additional configurations that are packaged as JSON files in the `configs/resources/config` folder. Config python code # Pipeline settings Source: https://docs.prophecy.ai/data-engineering/development/pipelines/pipeline-settings Control how your pipeline runs Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Review the various settings available for each pipeline, including Spark settings, code customization, development preferences, job sampling, run settings, and initialization code. Pipeline settings ## Spark
| Setting | Description | | -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Spark version | The Spark version associated with the pipeline. | | Mode | The pipeline's mode (either batch or streaming). | | Spark configuration | Name-value pairs will be set inside the Spark runtime configurations as `spark.conf.set(name, value)`.

You can edit the JSON directly by toggling the **Edit as JSON** button. | | Hadoop configuration | Name-value pairs will be set inside the Hadoop configuration as `spark.sparkContext.hadoopConfiguration.set(name, value)`.

You can edit the JSON directly by toggling the **Edit as JSON** button. |
## Code
| Setting | Description | | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Package Name | The name of the package if the project is published to the [Package Hub](/data-engineering/extensibility/package-hub/package-hub). | | Config Package Name | A unique name for the pipeline's configuration package.
Only pipelines made before Prophecy 3.4.5.0 may need to have a custom config package name. | | Custom Application Name | The name of the Spark job that appears in the Spark interface. | | Allow Configuration Updates (Scala only) | When enabled, you can override configuration values using a script.
For example, if you add a Script gem to the pipeline, you can write something like `Config.current_date_var = "2024"` to set the value of that variable. | | Enable pipeline monitoring | The option to turn pipeline monitoring on or off. | | Enable jinja based configuration | The option to turn [jinja syntax](/data-engineering/development/pipelines/configuration#use-configuration-variables) on or off. |
## Development
| Setting | Description | | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Visual Language | The programming language (SQL, Scala, or Python) used for expressions inside of gems. If you change the visual language while developing your pipeline, Prophecy will automatically convert expressions into the chosen language. The [Expression Builder](/data-engineering/gems/expression-builder) will adapt to the language as well. |
## Job
| Setting | Description | | ---------------------- | ------------------------------------------------------------ | | Job Data Sampling | A toggle to enable or disable data sampling during job runs. | | Job Data Sampling Mode | The sampling mode used during job runs. |
## Run Settings
| Property | Description | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Limit Input Records | When enabled, this limits the number of rows being operated on, makes development time faster, and reduces computation cost. Depending on how your pipeline is constructed, you might run into some issues when limiting records. If the number of records is too small, you might accidentally exclude records that, for example, match a join condition. This would result in an empty output. | | Data Sampling | Data sampling is enabled by default so you can view interim data samples while developing your pipeline. Learn about different [sampling modes](/data-engineering/development/runs/data-sampling). | | Configuration | This setting determines which [configuration](/data-engineering/development/pipelines/configuration) will be used during a pipeline run. |
## Initialization Code
| Setting | Description | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | UDF initialization code | The code that is run before initializing any UDF method. In this field, you can define variables, add common classes, include common imports, and more. |
# Pipelines for Data Engineering Source: https://docs.prophecy.ai/data-engineering/development/pipelines/pipelines Flows that represent the data journey Applicable to the [Enterprise Edition](/data-engineering/administration/platform/editions) only. Pipelines are groups of data transformations that you can build from a **visual** or **code** interface. When using the visual interface, each component of a pipeline is automatically compiled into code that you can reuse and customize. Under the hood, pipelines are based on Spark-native code. Pipelines are ideal for Spark environments like Databricks or EMR, particularly for tasks such as complex data ingestion (e.g., loading data from Salesforce or JDBC), handling advanced data transformations (e.g., working with complex data types), and supporting machine learning workflows. ## Creation If you want to create a new pipeline, you can do so from the **Create Entity** page in the left sidebar. You can also create pipelines directly within the project editor. The following table describes the parameters for pipeline creation. | Field | Description | | ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | | Project | The project to create the pipeline in. This controls access to the pipeline, groups pipelines together, and lets you use datasets in the project. | | Branch | The Git branch to use for pipeline development. | | Name | The name of the pipeline. | | Mode | Whether the pipeline will be batch mode (collect and process data in scheduled intervals) or streaming (ingest and transmit data in real-time). | | Description | A field to describe the purpose of the pipeline. | ## Project editor When building your pipelines, it helps to be familiar with the project editor interface. The following table describes different areas of the project editor. | Callout | Component | Description | | ------- | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 1 | Project tab | A list in the left sidebar that shows all of the project components. When in code view, this project tab shows the file directory with the code components. | | 2 | Environment browser | A list in the left sidebar that lets you browse different assets in your connected execution environment. For example, you can browse the Unity Catalog if you are attached to a Databricks fabric. | | 3 | Search | A search that lets you find different components like gems in your project and pipelines. | | 4 | Canvas | The area in the center of the project where you build your pipelines visually. | | 5 | Header | A menu that includes various configurations such as project settings, dependency management, cluster attachment, scheduling, and more. It also provides a toggle to switch between the Visual view and Code view. | | 6 | Footer | A menu that includes diagnostic information, execution metrics, execution code, and the Git workflow. | See these components marked in the image below. Project Editor ### Canvas Let's take a closer look at the pipeline canvas. The canvas includes: * **Canvas**: space to add and connect gems. * **Gem drawer**: toolbox that contains all available gems. * **Run button**: click to [execute the pipeline interactively](/data-engineering/development/runs/execution). * **Copilot**: AI assistant to help build your pipeline. Pipeline canvas ## Metadata To view a list of pipelines in Prophecy, navigate to the **Metadata** page from the left sidebar. For more granular metadata, click into a pipeline. Pipeline metadata can also be accessed from the header of the project editor. Pipeline metadata The table below describes the different tabs inside an individual pipeline's metadata. | Tab | Description | | --------- | ------------------------------------------------------------------------------------------------------------------ | | Info | A list of the input and output datasets of the pipeline. You can also edit the pipeline name and description here. | | Relations | A list of jobs and subgraphs that include the pipeline. | | Code | The code that is stored in the Git repository for the pipeline. | | Runs | A history of pipeline runs per fabric. | # Secrets in pipeline configurations Source: https://docs.prophecy.ai/data-engineering/development/pipelines/secrets-configs Store secrets in pipeline config Applicable to the [Enterprise Edition](/data-engineering/administration/platform/editions) only. Secrets provide secure authentication for connecting to various data tools like Salesforce, REST APIs, and Snowflake by storing sensitive credentials in centralized secret providers rather than exposing them in code. Any time you need to enter credentials in a Prophecy pipeline, you will be prompted to insert a secret or insert a configuration. Follow this guide to understand how to use configurations for secrets. ## Use cases There are a few cases in which you might want to add secrets to pipeline configurations, rather than inserting the secrets directly into gems. * You want different credentials to be used for different pipeline runs (for example, development versus production runs). * You use the secret in multiple gems in the pipeline. If your secret is in a pipeline config, then you can change the value once in the config and it will apply to all gems. ## Example In this example, we demonstrate using Databricks secrets to configure Snowflake credentials to establish a connection to Snowflake within a gem. 1. **Create secrets in Databricks.** [Create your secret scope and keys in Databricks](https://docs.databricks.com/security/secrets/index.html). For this example, create a secret for your Snowflake username and a secret for the password. Assume we created scope `demo-scope` and added two secrets with key `snowflake-username` and `snowflake-password`. 2. **Create pipeline config to map to secrets.** Add configs of Type `databricks_secret` in [Pipeline Configs](/data-engineering/development/pipelines/configuration). Let's say we call it `snowflake_user` and `snowflake_pass`. Open pipeline configuration 3. **Provide values to the config created.** Now, lets add value for the created configs `snowflake_user` and `snowflake_pass` in the default config. You can also add multiple values in different configs. For value, add the scope and key you created for your secret in the first step and save it. It's now ready to be used in your gems. Add default values 4. **Add a Snowflake gem to your pipeline and use the config.** Use the Config with syntax as `${snowflake_user}` and `${snowflake_pass}` in the username and password field respectively and define all other required fields in the gem as is. Use in a Snowflake gem # Conditional execution Source: https://docs.prophecy.ai/data-engineering/development/runs/conditional-execution Conditionally run or skip transformations within pipelines Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. For granular data processing control, you can conditionally run or skip transformations within gems in your pipeline. This means you can configure **pass-through conditions** on gems to dynamically control whether a transformation is executed. You also have the option to configure a **removal condition** on a gem, which not only skips the transformation but also removes the gem and all associated downstream transformations from pipeline execution. ## Configure conditions To configure a condition on a gem: 1. Click the **...** (ellipsis) on a gem. 2. Select the **Add Condition** option. 3. Choose the **Pass through condition** or **Remove condition** option. 4. Write your condition in Scala or Python, depending on your project language. Add a condition When a condition is set on a gem, a (C) symbol will appear before the gem name. When a gem meets a pass-through or removed condition, the interims will not be displayed on the edges associated with that gem. ## Pass-through condition Pass-through conditions let you skip the transformation of a gem or subgraph and maintain the input data as the output data. This ensures that the data passes through the gem or subgraph without any modification. When using pass-through conditions, be aware that: * The gem or subgraph must be connected in the pipeline, meaning it should have both an input port and an output port. This allows the data to flow through the gem. * Pass-through conditions are not applicable to source and target elements within the pipeline. These elements represent the data source and destination and do not involve any transformation logic. ## Removal condition When a removal condition is met, the associated gem and all of its downstream gems are excluded from the pipeline execution. Unlike pass-through conditions, you can use removal conditions on Source and Target gems. ### Example: Run if input record count > 0 Assume you only want to write to a target table if the input record count is greater than zero. You can use a removal condition to do so: 1. Click the **...** (ellipsis) on the Target gem. 2. Select the **Add Condition** option. 3. Note that you can only define a removal condition. 4. Use the condition `df_.count() > 0`. This example is Python code. 5. Click **OK** to save. Remove condition on Target gem # Data sampling Source: https://docs.prophecy.ai/data-engineering/development/runs/data-sampling Choose when to sample data during interactive execution Available for [Enterprise Edition](/data-engineering/administration/platform/editions) only. Prophecy gives you control over when and where data samples are generated during interactive pipeline execution. This helps you optimize for speed, visibility, or compatibility with your fabric setup. You can customize data sampling at three levels: * **Gem level**: Turn sampling on or off for individual or multiple gems. * **Pipeline level**: Set the sampling mode used during interactive runs. * **Fabric level**: Enable or disable data sampling for pipelines running on particular fabrics. This page describes how to set up and use data samples for your use cases. ## Interactive run configuration You can adjust data sampling settings for each pipeline through the **Interactive Run Configuration** panel. You can also access the same settings in [Pipeline Settings](/data-engineering/development/pipelines/pipeline-settings#run-settings). 1. Hover the large **play** button in the canvas. 2. Click on the **ellipses** that appears on hover. 3. Toggle data sampling on or off. 4. When data sampling is on, select your [preferred mode](#data-sampling-modes) from the **Data Sampling** dropdown. Interactive run configuration ## Data sampling modes Prophecy provides the following data sampling modes. | Mode | Samples generated | Use case | | ----------------- | ------------------------------------------------------------------------- | ---------------------------------- | | **All** (default) | After every gem, excluding Target gems. | Full visibility | | **Selective** | When **Data Preview** enabled per gem. [Learn more](#selective-sampling). | Full control per gem | | **Sources** | Only after Source gems. | Focus on inputs | | **Targets** | Only before Target gems. | Focus on outputs | | **IO** | Only after Sources and before Targets (not between intermediate gems). | High-level input/output inspection | ### Selective sampling Selective data sampling gives you granular control by letting you enable or disable data samples for individual gems. To control data sampling for gems: * **Single gem**: Select the Data Preview checkbox in the gem's [action menu](/data-engineering/gems/gems). * **Multiple gems**: Select multiple gems by dragging, then click the Data Preview button in the bottom menu. When Data Preview is disabled for a gem, its output appears pale after pipeline execution, indicating no sample was generated. Click the pale output to load the data sample on demand. The output icon will then display in normal bold colors. Selective Prophecy recommends using selective sampling mode for all users, regardless of your Spark provider. Selectively-generated samples load up to 10,000 rows (or 2 MB payload) by default. Set the following environment variables for your Spark cluster to modify this behavior: * `EXECUTION_DATA_SAMPLE_LOADER_MAX_ROWS`: Max number of rows (default is 10,000 rows). * `EXECUTION_DATA_SAMPLE_LOADER_PAYLOAD_SIZE_LIMIT`: Max payload size (default 2 MB). * `EXECUTION_DATA_SAMPLE_LOADER_CHAR_LIMIT`: Per column character limit (default 200 KB). Values exceeding the limit are truncated. ## Cached interims When you change data sampling settings and re-run a pipeline, some data samples may appear grayed out. These cached samples are from previous runs and may not reflect your current data or pipeline changes. Cached interims ## Record counts In addition to data samples, Prophecy can also display the total record count of datasets between gems. This works for **selective data sampling** mode only. To display the record count for a certain gem output: 1. Click the **...** (ellipsis) on the gem. 2. Select the **Record Count** checkbox. Gem menu with Record Count checkbox You can enable the record count and leave the data preview option disabled. They are independent of each other. The following image shows a pipeline with the record count enabled on the DataCleansing and Reformat gems. Notice that the Reformat gem is the only gem that has data preview enabled. Pipeline with record count enabled ## Fabric settings In a fabric, you can enable or disable data sampling and override pipeline-level settings when a pipeline runs on that fabric. You can access this option in the **Advanced** tab of a fabric. A common use case is preventing sample data generation in **production** pipelines. Create a new model test By default, only team admins can access the Advanced tab in a fabric. However, there are two flags you can set in your deployment to change this behavior: * `ALLOW_FABRIC_ACCESS_CLUSTER_ADMIN`: Grants cluster admins full access to fabrics, even if they are not team admins. * `DISALLOW_FABRIC_CODEDEPS_UPDATE_TEAM_ADMIN`: Prevents team admins from modifying the data sampling settings within a fabric. # Run types Source: https://docs.prophecy.ai/data-engineering/development/runs/execution Different ways you can run Prophecy pipelines Applicable to the [Enterprise Edition](/data-engineering/administration/platform/editions) only. Prophecy executes pipelines based on the type of Spark cluster and the operation being performed. You can run pipelines interactively, schedule them, or execute them using the built-in Spark shell. Each execution provides information through logs and metrics to help you manage and monitor data transformations. ## Interactive execution When you run a pipeline in the pipeline canvas, Prophecy generates interim **data samples** that let you preview the output of your data transformations. There are two ways to run a pipeline interactively: * **Click the large play button on bottom of the pipeline canvas.** The whole pipeline runs. * **Click the play button on a gem.** All gems up to and including that gem run. This is a partial pipeline run. Interactive run options After you run your pipeline interactively in the canvas, data samples will appear between gems. These previews are temporarily cached. Learn about how these [data samples](/data-engineering/development/runs/data-sampling) are generated or discover the [Data Explorer](/data-engineering/development/data-explorer/data-explorer). ## Gem execution order When pipelines have multiple branches, gems execute sequentially, rather than running in parallel. You can visualize this sequence by switching to the **Code** view of your pipeline, where operations appear in their exact execution order. ```python theme={null} def pipeline(spark: SparkSession) -> None: df_Orders = Orders(spark) df_Customers = Customers(spark) df_By_CustomerId = By_CustomerId(spark, df_Orders, df_Customers) df_Cleanup = Cleanup(spark, df_By_CustomerId) df_Sum_Amounts = Sum_Amounts(spark, df_Cleanup) Customer_Orders(spark, df_Sum_Amounts) ``` In this example, each step executes completely before the next one begins: * Step 1 (Orders) runs and completes entirely * Step 2 (Customers) runs only after Orders finishes * Step 3 (Join) waits for both previous steps to be complete * Steps 4, 5, and 6 continue this sequential pattern The order that you see in the Code view can be changed by manually setting [gem phases](/data-engineering/gems/gems#gem-phase). Gems running on Spark are limited to sequential execution. Conversely, when gems run in a SQL warehouse, Prophecy leverages parallel execution when dependencies allow. ## Scheduled execution When you create [jobs](/data-engineering/orchestration/databricks-jobs) in Prophecy, you schedule when certain pipelines will run. Prophecy executes pipelines based on the [fabric](/data-engineering/fabrics/spark-provider/databricks/databricks) defined in the pipeline's job settings. Prophecy will automatically run jobs once relevant projects are released and deployed. ## Shell execution Prophecy comes with an built-in interactive Spark shell that supports both Python and Scala. The shell is an easy way to quickly analyze data or test Spark commands. The Spark context and session are available within the shell as variables `sc` and `spark` respectively. Interactive execution ## Execution information Once you run a pipeline, there are several ways for you to better understand the execution. | Callout | Information | Description | | ------- | ---------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | | **1** | Problems | Errors from your pipeline execution that will be shown in a dialog window, as well as in the canvas footer. | | **2** | Runtime logs | The progress with timestamps of your pipeline runs and any errors. | | **3** | Execution code | The code Prophecy runs to execute your pipeline. You can copy and paste this code elsewhere for debugging. | | **4** | Runtime metrics | Various Spark metrics collected during runtime. | | **5** | [Execution metrics](/data-engineering/fabrics/execution-metrics) | Metrics that can be found in the **Metadata** of a pipeline, or from the **Run History** button under the **...** menu. | Use the image below to help you find the relevant information. Execution information ## Execution on Databricks Databricks clusters come with various [Access Modes](https://docs.databricks.com/clusters/create-cluster.html#what-is-cluster-access-mode). To use Unity Catalog Shared clusters, check for feature support [here](/data-engineering/fabrics/spark-provider/databricks/UCShared). When using `High Concurrency` or `Shared Mode` Databricks Clusters you may notice a delay when running the first command, or when your cluster is scaling up to meet demand. This delay is due to Prophecy and pipeline dependencies (Maven or Python packages) being installed. For the best performance, it is recommended that you cache packages in an Artifactory or on DBFS. Please [contact us](https://help.prophecy.io/support/tickets/new) to learn more about this. # Spark structured streaming Source: https://docs.prophecy.ai/data-engineering/development/spark-streaming/spark-streaming Learn about streaming data running on Spark Structured Streaming Prophecy no longer provides support for streaming pipelines. Please switch to batch pipelines for continued support. Prophecy provides native support for streaming data running on Spark Structured Streaming. This documentation assumes you are already familiar with how Structured Streaming works. For more information, you can consult the Structured Streaming documentation [here](https://spark.apache.org/docs/latest/structured-streaming-programming-guide.html). Streaming [pipelines](/data-engineering/development/pipelines/pipelines) work differently from batch pipelines: 1. Streaming applications are always running, continuously processing incoming data. 2. Data is processed in micro-batches, with the notable exception of [Continuous Triggers](https://spark.apache.org/docs/latest/structured-streaming-programming-guide.html#continuous-processing) (an experimental feature available in Spark3.3). Continuous triggers are not supported by Prophecy. 3. Streaming applications handle transient data rather than maintain the entire data. Aggregations and joins require watermarking for maintaining a limited state. 4. All Streaming datasets can behave similarly to Batch datasets using the Spark [`ForEachBatch`](https://spark.apache.org/docs/latest/api/python/reference/pyspark.ss/api/pyspark.sql.streaming.DataStreamWriter.foreachBatch.html), though `ForEachBatch` is not supported by Prophecy. The streaming capability is available for `Python` projects that do not use UC standard clusters. | Project Type | Spark Cluster Access Mode | Spark Cluster Type | Structured Streaming Capability | | ------------ | -------------------------------- | ------------------ | --------------------------------- | | Python | Dedicated (formerly single user) | UC, legacy | Supported as of Prophecy3.4.x | | Python | Standard (formerly shared) | UC | Not supported as of Prophecy3.4.x | | Scala | Any Mode | Any type | Not supported as of Prophecy3.4.x | ## Spark Structured Streaming using Prophecy IDE How to Create a Streaming pipeline Within a Prophecy `Python` project, a user can create a Structured Streaming pipeline using the Streaming(beta) mode. ### Working with a streaming pipeline To create a streaming pipeline, users can follow a process similar to creating a Batch pipeline in a `Python` project. Streaming pipelines work differently from Batch pipelines in the following ways: 1. Partial runs are not supported for streaming applications. A partial run is only allowed on a `Streaming Target` gem. 2. Streaming pipelines are long-running tasks and process data at intervals. Currently, they do not capture cumulative statistics. 3. Streaming pipelines are continuous and do not stop running. To terminate a Streaming pipeline, users need to click the "X" button. A Streaming pipeline is an ongoing process and will not terminate itself. 4. To deploy the pipeline on Databricks, users can follow the same process described [here](/data-engineering/orchestration/databricks-jobs). A scheduled job will check if the Streaming pipeline is running every X minutes. If the pipeline is not running, the job will attempt to start it. ### Streaming Sources and Targets Spark Structured Streaming applications have a variety of source and target components available to construct Piplines. Streaming source gems render to `spark.readStream()` on the Spark side. Currently, we support file stream-based sources and targets, warehouse-based targets, and event stream-based sources and targets. Additionally, any batch data sources can be used in a streaming application. Batch data sources are read using the `spark.read()` function at every processing trigger (due to Spark evaluating lazily). More on triggers [here](https://spark.apache.org/docs/latest/structured-streaming-programming-guide.html#triggers). For more information on Batch Source and Target gems, click [here](/data-engineering/gems/source-target). ### Streaming Transformations For more information on Streaming Transformations, click [here](./streaming-transformations.md). # Event-based Source: https://docs.prophecy.ai/data-engineering/development/spark-streaming/streaming-sources-and-targets/streaming-event-gem Event-based Source and Target Gems for Streaming Data Applications Prophecy no longer provides support for streaming pipelines. Please switch to batch pipelines for continued support. ## Event-based Sources and Targets Prophecy supports **Kafka Streaming** Source and Target. More information on supported Kafka Source and Target options are available [here](https://spark.apache.org/docs/latest/structured-streaming-kafka-integration.html). The Kafka gem allows inferring the schema of the events by automatically populating the `value` column. Schema inference works with both JSON and AVRO file formats. A user is required to provide an example event for schema inference. ## Create a Kafka Source gem A Kafka Source gem allows the Streaming pipeline continuously pull data from a Kafka topic. The following options are supported: | **Property** | Optional | **Default Value** | **Comment** | | --------------------- | -------- | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | | Broker List | False | N/A | List of Kafka brokers separated by commas. For eg. `kdj-ibg1.us-east-2.aws.cloud:9092, kdj-ibg2.us-east-2.aws.cloud:9092,kdj-ibg3.us-east-2.aws.cloud:9092` | | **Group ID** | True | None | Consumer group ID. | | **Session Timeout** | False | 6000 | Corresponds to the `session.timeout.ms` field | | **Security Protocol** | False | SASL\_SSL | Supported values are `SASL_SSL`, `PLAINTEXT`, `SSL`, `SSL_PLAINTEXT` | | **SASL Mechanisms** | False | SCRAM-SHA-256 | SASL mechanism to handle username/password authentication. Supported values are `PLAIN`, `SCRAM-SHA-256` and `SCRAM-SHA-512`, `GSSAPI`, `OAUTHBEARER` | | **Kafka Topic** | False | N/A | Name of Kafka Topic to Consume | ### Entering Authentication Credentials * **Databricks Secrets (recommended)**: Use Databricks to manage your credentials * **UserName, Password**: Use **ONLY** for test deployments and during development. This writes credentials to Git repository, which isn't good practice. # File-based Source: https://docs.prophecy.ai/data-engineering/development/spark-streaming/streaming-sources-and-targets/streaming-file-gem File-based Source and Target gems for Streaming Data Applications Prophecy no longer provides support for streaming pipelines. Please switch to batch pipelines for continued support. ## File-based Streaming Sources and Targets For file stream sources, incoming data files are incrementally and efficiently processed as they arrive in cloud storage. No additional setup is necessary, and cloud storage only needs to be accessible from the User's fabric. Autoloader is available for use with a Databricks fabric and supports loading data directory listing, as well as using file notifications via AWS's Simple Queue Service (SQS). More on Autoloader [here](https://docs.databricks.com/ingestion/auto-loader/index.html). For different Cloud Storages supported by Autoloader, please check [this](https://docs.databricks.com/ingestion/auto-loader/file-detection-modes.html) page. When you select Format and click NEXT, this Location Dialog opens: File Streaming ## Databricks Auto Loader Databricks fabrics can utilize [Auto Loader](https://docs.databricks.com/ingestion/auto-loader/index.html). Auto Loader supports loading data directory listing as well as using AWS's Simple Queue Service (SQS) file notifications. More on this [here](https://docs.databricks.com/ingestion/auto-loader/file-detection-modes.html). Stream sources using Auto Loader allow [configurable properties](https://docs.databricks.com/ingestion/auto-loader/options.html#file-format-options) that can be configured using the Field Picker on the gem: Autoloader Directory Listing Mode Autoloader Filer Notifiction Mode ## Formats Supported The following file formats are supported. The gem properties are accessible under the Properties Tab by clicking on `+` : 1. JSON: Native Connector Docs for Source [here](https://spark.apache.org/docs/3.5.8/api/python/reference/pyspark.ss/api/pyspark.sql.streaming.DataStreamReader.json.html). Additional Autoloader Options [here](https://docs.databricks.com/aws/en/ingestion/cloud-object-storage/auto-loader/options#json-options). 2. CSV: Native Connector Docs for Source [here](https://spark.apache.org/docs/3.5.7/api/python/reference/pyspark.ss/api/pyspark.sql.streaming.DataStreamReader.csv.html). Additional Autoloader Options [here](https://docs.databricks.com/aws/en/ingestion/cloud-object-storage/auto-loader/options#csv-options). 3. Parquet: Native Connector Docs for Source [here](https://spark.apache.org/docs/3.5.8/api/python/reference/pyspark.ss/api/pyspark.sql.streaming.DataStreamReader.parquet.html). Additional Autoloader Options [here](https://docs.databricks.com/aws/en/ingestion/cloud-object-storage/auto-loader/options#parquet-options). 4. ORC: Native Connector Docs for Source [here](https://spark.apache.org/docs/3.5.7/api/python/reference/pyspark.ss/api/pyspark.sql.streaming.DataStreamReader.orc.html).Additional Autoloader Options [here](https://docs.databricks.com/ingestion/auto-loader/options.html#orc-options). 5. Delta: A quickstart on Delta Lake Stream Reading and Writing is available [here](https://docs.databricks.com/structured-streaming/delta-lake.html#delta-table-as-a-source). Connector Docs are available [here](https://docs.delta.io/latest/delta-streaming.html). Note, that this would require installing the Spark Delta Lake Connector if the user has an on prem deployment. We have additionally provided support for Merge in the Delta Lake Write Connector. (uses `forEatchBatch` behind the scenes). ## File-based Streaming Tutorial