AI penetration testing for Australian businesses.
Manual penetration testing of chatbots, copilots, retrieval pipelines and agents with tool access, under the same written scope and rules of engagement as any other system, extended to prompt injection, jailbreak attempts and data leakage through retrieval.
Black Shard conducts AI penetration testing: manual testing of chatbots, copilots, retrieval-augmented systems and agents with access to internal tools. The same written scope, rules of engagement and reproduced findings that apply to a standard penetration test apply here, extended to the AI-specific surface: prompt injection, jailbreak and guardrail bypass, insecure output handling, data leakage through retrieval, and tool and agent permission abuse.
Black Shard is an Australian software engineering and cybersecurity firm, headquartered in Brisbane, and builds and operates production AI systems as well as testing them. Testing runs against systems wherever in Australia they are operated. It requires a defined scope, provisioned access and an agreed testing window.
Scoping an AI penetration test
Every AI penetration test starts with a written scope and quote, the same as any other engagement. Scope is set against what the AI feature does and what it touches: a chatbot or copilot answering from a knowledge base, a retrieval pipeline pulling from client data, an agent that can call internal tools, or a model-facing API consumed by another system. The quote names its drivers: the number of AI features and integration points in scope, whether retrieval or tool access is involved, tenancy count for multi-tenant products, and whether any of the work needs to run on-site.
Rules of engagement are agreed in writing before testing begins, extended to cover the AI-specific surface: the prompts and inputs that are in scope, any content the system is excluded from generating even to prove a finding, and a stop-work procedure if testing surfaces a live production issue. Access granted for the engagement is least-privilege and time-boxed to the testing window, and covers what a normal user of the feature would have plus whatever the scope calls for on the configuration behind it, such as the system prompt or guardrail settings.
What makes testing an AI system different
A chatbot, copilot or agent sits on top of the same authentication, session handling and access control as the rest of the application, and that layer is tested the way it always is. The AI feature adds a surface that standard web and API testing does not reach: natural language is itself an input an attacker can use to redirect the system's behaviour, and the system's output can be trusted downstream in ways a fixed API response is not.
Model behaviour is also probabilistic. A prompt injection or jailbreak attempt that fails once can succeed on a later attempt with the same input, so testing runs multiple attempts against each technique instead of treating one failed attempt as a clean result. A guardrail, whether a system prompt, a content filter or a permission boundary on a connected tool, is tested as a control that can fail rather than assumed to hold.
What we test
LLM penetration testing is run against the OWASP LLM Top 10 for the AI-specific surface, alongside the standard OWASP guidance already applied to the application the feature sits inside.
- Prompt injection, direct and indirect: an instruction typed straight into the input, or planted in a document, email or web page the system later retrieves and treats as trustworthy
- Jailbreak and guardrail bypass: attempts to make the system act outside the behaviour boundary it was configured with
- Insecure output handling: whether the system's output is trusted by whatever reads it next, including being rendered as executable content or passed to another system unchecked
- Data leakage through retrieval: whether a retrieval-augmented system can be made to surface data outside the requester's access, including across tenants or permission levels
- Tool and agent permission abuse: whether an agent's access to internal tools and systems can be escalated, redirected or used outside its intended scope
- Tenant isolation of AI features: for multi-tenant products, whether the AI feature enforces the same tenant boundary as the rest of the system
- Model-facing API abuse: rate limiting, cost abuse and unauthorised access to the API surface the model sits behind
Findings and reporting
Every finding in the report has been reproduced by a tester before it is written up, with repeated attempts recorded against each AI-specific technique given the probabilistic nature of model behaviour. Automated tooling is used for coverage; its output is reported as a finding only after a person has confirmed it. Each finding is written with severity, business impact, reproduction steps and the technical detail required to remediate it, whether the fix sits in the system prompt, the guardrail configuration, the retrieval permission model or the surrounding application.
Findings are the client's information. They are held in confidence, and any disclosure to a third-party vendor, including a model provider, is made only with the client's consent.
- A written scope, quote and rules of engagement before testing starts
- A findings report with severity, business impact, reproduction steps and remediation detail for each issue
- A debrief with the tester by video call, or in person in Brisbane
- A re-test of agreed fixes, limited to the findings in the original report
Built by engineers who run production AI
Black Shard builds and operates production AI systems as well as testing them. Aurii, clinical software built and operated by Black Shard, runs Azure AI Speech recognition and OCR pipelines against clinical dictation and documents, in a multi-tenant Postgres database enforced with row-level security in Azure's Australia East region. Black Shard's AI and automation engineering practice builds speech-to-text and OCR pipelines and workflow automation, with an audit trail and guardrails around each automated action.
An AI penetration test run by a firm that builds and operates this class of system tests the guardrail, the tenant boundary and the audit trail against the engineering decisions behind them.
How much does AI penetration testing cost?
AI penetration testing is quoted as a fixed scope before work starts. Cost is driven by the number of AI features and integration points in scope, whether the system is a chatbot or copilot over a knowledge base, a retrieval pipeline, or an agent with tool access, tenancy count for multi-tenant products, and whether on-site work is included. The re-test sits inside the fixed scope.
Send a brief to [email protected] and scoping starts from the detail you have.
Questions, answered
- Is this the same as your standard penetration testing?
- It runs under the same practice: written scope, rules of engagement, findings reproduced before reporting, and a re-test of agreed fixes. The AI feature adds a surface on top: prompt injection, jailbreak attempts, insecure output handling and data leakage through retrieval.
- Do you test for prompt injection?
- Yes. Prompt injection testing covers direct injection, where an attacker types the instruction straight into the input, and indirect injection, where the instruction is planted in a document, email or web page the system later retrieves and treats as trustworthy.
- What is chatbot security testing?
- Chatbot security testing examines a customer-facing or internal chatbot for prompt injection, jailbreak and guardrail bypass, and whether a conversation can be steered into retrieving or disclosing data outside the requester's access.
- What methodology do you test AI systems against?
- The OWASP LLM Top 10 for the AI-specific surface, and OWASP's standard web and API guidance for everything the AI feature sits inside: authentication, session handling and access control.
- Can you test an AI agent that has access to internal tools or systems?
- Yes. Where an AI feature can call tools, trigger workflows or take actions on other systems, testing covers whether that access can be escalated, redirected or used outside its intended scope.
- Is a re-test included?
- A re-test of the findings in the original report can be included in the fixed scope. Model behaviour is probabilistic, so a re-test confirms the fix holds across repeated attempts rather than a single pass.
Scope an AI penetration test with Black Shard.
Brisbane head office. Work delivered across Australia.