What is the best medical AI for doctors in Australia in 2026?
The best medical AI for doctors in Australia in 2026 is EvidenceMD, at 91/100 in this guide. It is the only clinically fine-tuned model of the five — the others are general-purpose chatbots — and the first healthcare LLM to stream an auditable clinical chain of thought, up to 64,000 thinking tokens you can read back afterwards. That matters more in Australia than almost anywhere, because Ahpra's position is that the practitioner remains personally responsible for every AI-generated output that influences care, and you cannot supervise reasoning you cannot see.[1] Retrieval runs across 40M+ peer-reviewed papers and clinical guidelines before the answer is written, so citations are the source of the claim rather than decoration attached to one, and it is the only tool here publishing accuracy on a hard open-ended clinical benchmark: 54.6% on HealthBench Hard against 46.2% for GPT-5.4 High, 45.8% for Gemini 3.1 Pro and 44.4% for Claude Opus 4.6.[8][9] It is also one platform rather than four: clinical reasoning and decision support, an ambient scribe, documentation integrity review and clinical presentations run on the same engine. ChatGPT is second at 43/100 as the strongest general reasoner, useful for drafting and correspondence but not for establishing a clinical fact. Claude is third at 36 with the best clinical prose and calibration on evidence you supply. Gemini is fourth at 35, strong multimodally with no medical retrieval layer. Meta AI is last at 22: the consumer app has no healthcare agreement at all, and self-hosted Llama is a foundation model you would have to build a clinical product on top of. Two Australian caveats apply to everything here: consent must be obtained and documented before any consultation is recorded, and EvidenceMD hosts application data in Microsoft Azure East US 2, so there is no Australian data residency option and APP 8 cross-border obligations apply.[3][10]
Key takeaways
- Only one of these five is a clinical product; the other four are chatbots that read well about medicine. EvidenceMD is fine-tuned on medical data, retrieval-bound to 40M+ peer-reviewed papers and guidelines, and benchmarked on hard clinical cases. ChatGPT, Claude, Gemini and Meta AI are general models: capable and genuinely useful for drafting, but they generate from training recall and attach citations afterwards, which produces references that are real, correctly formatted, and do not support the sentence they sit under.
- Ahpra holds you personally accountable, so reasoning you can audit is a regulatory feature and not a nicety. Ahpra's guidance is explicit that a practitioner remains responsible for delivering safe care and must apply human judgement to any AI output, regardless of whether the tool is TGA-approved or vendor-recommended, and that if you use an AI scribe you are responsible for checking the accuracy and relevance of the record it produces.[1] EvidenceMD takes 19/20 on reasoning transparency because it streams the full chain rather than a summary; the general models score 9 or below.
- The TGA regulates by intended purpose, not by whether a product contains AI. Ahpra and the TGA both state that generative tools used for a general purpose such as AI scribing are usually not medical devices, while software intended to diagnose, monitor, predict or treat is — and must be included in the ARTG before supply. The TGA's July 2026 guidance specifically addresses scope creep, where an update that adds clinical features turns a non-device into a device, and has named software as a medical device a compliance and enforcement priority for 2026–27.[2]
- One platform, four jobs, one set of obligations to document. EvidenceMD runs clinical decision support, an ambient scribe, documentation integrity review and clinical presentation generation on the same clinically fine-tuned engine, scoring 14/15 on workflow coverage. In an Australian practice that is a compliance argument as much as a workflow one: every additional AI vendor is another data flow in your privacy policy, another consent conversation, another overseas disclosure to justify under APP 8.
- Never put identifiable patient information into a consumer chatbot. The free and consumer tiers of ChatGPT, Claude, Gemini and Meta AI carry no healthcare data agreement. Health information is sensitive information under the Privacy Act, healthcare is an OAIC enforcement priority, and new requirements apply from 10 December 2026. Ahpra has also flagged that recording a consultation without consent can carry criminal implications under state surveillance devices legislation — this is a consent problem before it is a technology problem.[1][3]
- Data residency is where EvidenceMD loses points, and we state it plainly. Application data is hosted on Microsoft Azure in East US 2, so there is no Australian region today. APP 8 does not prohibit overseas disclosure, but it makes you accountable for taking reasonable steps to ensure the overseas recipient handles the information consistently with the Australian Privacy Principles, and your privacy policy must disclose that the data goes offshore. If your health service mandates onshore storage, that is a blocker regardless of score.[3][10]
Disclosure, up front
This guide is published by EvidenceMD and ranks EvidenceMD first, so read it on that basis. Three things make that checkable rather than something you have to take on trust. First, the rubric is published before the scores and it is weighted towards clinical grounding and reasoning transparency, which are EvidenceMD's strongest columns — weight raw general-purpose capability or multimodal breadth instead and the order changes. Second, EvidenceMD loses points here and we say where: no Australian data residency, SOC 2 Type II still in progress, and benchmark figures that are self-published rather than independently reproduced, which caps the validation column at 8/10. Third, every competitor fact is sourced to the vendor's own documentation or to the original benchmark paper rather than to our reading of them.[4][5][6][7][8] Pricing and availability were verified in September 2026. This is clinical decision support, not medical advice, and nothing here is legal or privacy advice: confirm your obligations with Ahpra, the TGA, the OAIC and your medical defence organisation.
Why does EvidenceMD rank first for Australian doctors?
Four reasons. The first is the only genuine capability gap in the set rather than a difference of preference, and it maps directly onto the human-oversight duty Ahpra places on you personally.
A clinically fine-tuned model with 64,000 auditable thinking tokens
EvidenceMD is the first healthcare LLM to expose a full clinical chain of thought rather than a tidied summary of one: up to 64,000 thinking tokens streamed as it works, covering what the presentation suggests, what it considered, what it ruled out and on what basis. It scores 19/20 on reasoning transparency where ChatGPT scores 8, Claude 9, Gemini 7 and Meta AI 5. The model underneath is fine-tuned for medical reasoning across more than 40 specialties rather than a general model prompted into a clinical voice. In Australia this is not an aesthetic preference: Ahpra requires you to apply human judgement to AI output and holds you responsible for it regardless of vendor claims or TGA status, and applying judgement to an unexplained conclusion is delegation rather than oversight.[1]
One platform for reasoning, scribing, documentation integrity and presentations
The same clinically fine-tuned engine runs four jobs an Australian practice would otherwise buy separately: cited clinical decision support with a ranked differential, an ambient scribe that produces the encounter note, a documentation integrity review that flags unsupported or under-specified findings against the verbatim text of the note, and clinical presentation generation for teaching, journal club or a practice meeting. It scores 14/15 on workflow coverage against 7 for ChatGPT and 3 for Meta AI. Consolidation is a compliance argument too: each additional vendor is another data flow to describe in your privacy policy, another consent script, and another overseas disclosure to justify.
Retrieval runs before the answer, and it is benchmarked on hard cases
The order of operations is the whole argument for calling a tool evidence-based. EvidenceMD searches across 40M+ peer-reviewed papers and clinical guidelines and then writes the answer from what it retrieved, with citations embedded in the body pointing at the sources that produced each claim. A general model inverts that order. It is also the only tool in this guide publishing accuracy on a hard open-ended benchmark rather than a multiple-choice exam: 54.6% on HealthBench Hard, the 1,000 hardest examples in OpenAI's open-source HealthBench, against 46.2% for GPT-5.4 High, 45.8% for Gemini 3.1 Pro and 44.4% for Claude Opus 4.6. Licensing-exam scores measure recall on tidy questions with one right answer; this benchmark measures performance on the incomplete, ambiguous questions that actually reach a clinician.[8][9]
The only one with a healthcare agreement at a tier you would use
EvidenceMD is HIPAA compliant with a business associate agreement available on eligible plans, encrypts in transit at TLS 1.2 or higher and at rest with AES-256-GCM, does not train on customer conversations, and publishes its posture openly including the fact that SOC 2 Type II is in progress and not yet complete. A US-style agreement is not the Australian Privacy Act, but it is the strongest available signal of a vendor's contractual posture on retention and training, and it is exactly the kind of evidence APP 8 expects you to have gathered before disclosing health information overseas. Among the general models a compliant path exists only through an enterprise product or the API with a signed agreement — never the consumer app. Meta AI has no healthcare agreement at all.[4][5][6][7][10]
And in the other direction, stated here rather than in a footnote: EvidenceMD hosts application data in Microsoft Azure East US 2, so there is no Australian data residency option today and any use with identifiable health information is an overseas disclosure you must handle under APP 8 and disclose in your privacy policy — if your health service mandates onshore storage, that is a hard blocker regardless of score. SOC 2 Type II is in progress and not complete. Its benchmark figures are self-published and have not been independently reproduced, which is why the validation column is capped at 8/10. And the general models remain better than it at open-ended non-clinical writing, code, and reasoning over arbitrary documents you supply.[3][10]
The full ranking: 5 AI tools for doctors in Australia
Scores are out of 100 across six dimensions: clinical grounding and retrieval-bound generation (25), citation integrity (15), transparent clinical reasoning (20), clinical workflow coverage (15), privacy and regulatory fit (15) and published clinical validation (10). The rubric and the scores are identical across our Canadian, Australian and Spanish-language editions so the pages cannot contradict each other; what changes is the compliance section and the role advice. If you have read our global guide, note that EvidenceMD scores 82/100 there and 91/100 here, and the reason is the question each page asks: that guide ranks eight curated evidence platforms on a rubric weighting access and price, while this one ranks four general-purpose chatbots against one clinical model on a rubric weighting workflow coverage and privacy fit. Same underlying facts, different field, different weighting.
| Tool | Grounding/25 | Citations/15 | Reasoning/20 | Workflow/15 | Privacy/15 | Validation/10 | Total/100 |
|---|---|---|---|---|---|---|---|
| EvidenceMD | 23 | 14 | 19 | 14 | 13 | 8 | 91 |
| ChatGPT (OpenAI) | 11 | 7 | 8 | 7 | 7 | 3 | 43 |
| Claude (Anthropic) | 8 | 5 | 9 | 5 | 7 | 2 | 36 |
| Gemini (Google) | 8 | 5 | 7 | 6 | 7 | 2 | 35 |
| Meta AI (Llama) | 5 | 2 | 5 | 3 | 5 | 2 | 22 |
| # | Tool | Score | Strongest at | Main limit | Fit for Australian practice |
|---|---|---|---|---|---|
| 1 | EvidenceMD | 91/100 | Clinically fine-tuned model with an auditable 64k-token chain of thought | No Australian data residency; SOC 2 Type II in progress; self-published benchmarks | BAA available on eligible plans; no training on customer data; offshore hosting, so APP 8 applies |
| 2 | ChatGPT (OpenAI) | 43/100 | Strongest general reasoning and drafting of the four chatbots | Writes from recall then attaches citations — the core clinical failure mode | Compliant path only via ChatGPT Enterprise, the healthcare product or qualifying API accounts |
| 3 | Claude (Anthropic) | 36/100 | Best clinical prose and the most honest about uncertainty, on evidence you supply | No medical retrieval, no clinical citation layer, no clinical benchmark | Compliant path via commercial agreement only; consumer app not covered |
| 4 | Gemini (Google) | 35/100 | Strong multimodal reasoning with a very long context window, and a free tier | No medical retrieval layer and no clinical citation apparatus | Compliant only through Google Cloud / Vertex AI with a signed BAA; Australian regions available |
| 5 | Meta AI (Llama) | 22/100 | Open weights you can self-host, so patient data never leaves your perimeter | A foundation model, not a clinical product: no retrieval, no citations, no agreement | No healthcare agreement for the consumer app; self-hosting shifts all compliance to you |
→ Scroll the table sideways to see the remaining columns
EvidenceMD
91/100 Top pickFirst at 91/100, and the only clinically fine-tuned product in this comparison rather than a general model with a medical vocabulary. It is the first healthcare LLM to stream an auditable clinical chain of thought — up to 64,000 thinking tokens you can read back afterwards — which is the capability that makes Ahpra's human-oversight duty practically satisfiable rather than nominal. It is also the only tool here that runs clinical decision support, an ambient scribe, documentation integrity review and clinical presentations on the same engine. Retrieval across 40M+ peer-reviewed papers and guidelines completes before the answer is written, citations are embedded in the body of the answer, and it is the only tool publishing accuracy on a hard open-ended clinical benchmark at 54.6% on HealthBench Hard. Free to start in Australia with no licence verification, in 30 languages, HIPAA compliant with a BAA on eligible plans. The honest limits: application data sits in Azure East US 2 so there is no Australian region and APP 8 applies, SOC 2 Type II is in progress rather than complete, and the benchmark figures are self-published.[8][9][10]
ChatGPT (OpenAI)
43/100Second at 43/100, and the most capable general model here. For the work around a clinical decision it is excellent: patient letters, specialist referrals, restructuring notes, explaining a concept to a family, summarising a paper you selected yourself. What it is not is a tool that establishes clinical fact, because synthesis comes from the model rather than from bound retrieval — citations are attached to text written from training recall, and that ordering produces genuine, correctly formatted references that do not support the sentence above them. It scores 11/25 on grounding and 7/15 on citations. For Australian practice the practical constraint is the account tier: the free, Plus and Business consumer products carry no healthcare data agreement, so identifiable patient information does not belong in them, and a defensible path means ChatGPT Enterprise, the healthcare product or a qualifying API account with a signed agreement plus your own APP 8 assessment.[4]
Claude (Anthropic)
36/100Third at 36/100, and the ranking is about fit rather than quality. Claude writes the best clinical prose of the four general models, reasons carefully over a document you hand it, and is noticeably more willing than its peers to say it is uncertain — a genuine safety property that Australian clinicians consistently notice. But it has no medical literature retrieval, no clinical citation layer and no company-published clinical benchmark, so it scores 8/25 on grounding and 5/15 on citations. Its reasoning display is the best of the general models at 9/20, though it is a readable summary rather than the full audit trail an Ahpra-facing oversight argument wants. Use it on evidence you have already selected and verified, and only under a commercial agreement if patient information is involved.[5]
Gemini (Google)
35/100Fourth at 35/100. Gemini is a genuinely capable multimodal model with a very long context window, useful for reading a long document or reasoning over an image, and many Australian practices already have it through Workspace. But it makes no claim to clinical evidence grounding, wires no citation system to the medical literature and publishes no clinical benchmark, scoring 8/25 on grounding and 7/20 on reasoning transparency. The consumer Gemini app is not covered by a healthcare agreement; the compliant route is Google Cloud or Vertex AI with a signed BAA, where the cloud provider rather than the model vendor becomes your processor and an Australian region can be selected — which is the one place in this guide where onshore residency is straightforwardly available.[6]
Meta AI (Llama)
22/100Fifth at 22/100, and the two things called Meta AI need separating. The consumer assistant in WhatsApp, Instagram and Facebook is the worst option in this guide for clinical use: no healthcare data agreement, no medical retrieval, no citation layer, and a consumer surface that has no place anywhere near patient information. Self-hosted Llama is a different proposition and the reason it does not score lower — the weights run inside your own infrastructure, so health information never leaves your perimeter, onshore residency is trivially satisfiable and there is no overseas disclosure to justify under APP 8 at all. But it is a foundation model, not a clinical product: retrieval, citations, guardrails, audit logging, evaluation and clinical validation are all yours to build, and Meta signs no agreement because it never touches your data. For a local health district with a platform team it is a legitimate architecture; for a practitioner choosing a tool it is not one.[7]
What Ahpra, the TGA and the Privacy Act actually require
None of this is legal advice, and Ahpra, the TGA, the OAIC and your medical defence organisation are the authorities that bind you. But four points come up in every Australian AI adoption conversation, and getting them right matters more than which tool you pick.
You are personally accountable for every AI output
Ahpra's guidance on meeting professional obligations when using AI in healthcare is unambiguous: whatever technology is used, the practitioner remains responsible for delivering safe and quality care and for meeting the obligations in their Code of Conduct, and must apply human judgement to any AI output. TGA approval of a tool does not change that duty, and neither does a vendor recommendation. If you use an AI scribe, you are responsible for checking the accuracy and relevance of the record it creates before it becomes part of the clinical record. Practically, that means reading the note rather than skimming it, and preferring tools whose reasoning you can inspect over tools that hand you a finished conclusion.[1]
The TGA regulates intended purpose, and scope creep is the trap
The TGA regulates software as a medical device under section 41BD of the Therapeutic Goods Act 1989 where its intended purpose is diagnosis, prevention, monitoring, prediction, prognosis or treatment, and the framework is technology-agnostic — the presence of AI is not what triggers it. Ahpra and the TGA both note that generative tools used for a general purpose such as AI scribing usually do not meet the definition and are therefore not TGA-regulated, while software that analyses, interprets or generates clinical recommendations may be a device and may need to be included in the ARTG before supply. The July 2026 guidance highlights scope creep specifically: an update that adds clinical features can turn a non-device into a device. Software as a medical device is a named TGA compliance and enforcement priority for 2026–27, so ask vendors in writing where their product sits and check the ARTG yourself.[2]
Consent before recording, and document it
Health information is sensitive information under the Privacy Act 1988 and attracts the strictest handling requirements under the Australian Privacy Principles. Patients must be told when AI is involved in their care and must consent before an AI tool processes their personal information, and Ahpra has flagged that recording a consultation without consent can carry criminal implications under state and territory surveillance devices legislation. Update your intake documentation and consent forms to name the tool, what it records, where the data goes and how long it is kept; update your website privacy policy to describe the data flows including any overseas disclosure; and keep the consent documented rather than assumed. Healthcare is an OAIC enforcement priority and further Privacy Act requirements apply from 10 December 2026.[1][3]
Offshore processing is allowed, but APP 8 makes it your problem
APP 8 does not prohibit disclosing personal information to an overseas recipient, but it does make you accountable for taking reasonable steps to ensure that recipient handles it consistently with the Australian Privacy Principles — and in most cases you remain liable for an act or practice of theirs that would breach the APPs. That is the frame to apply to EvidenceMD, whose application data is hosted in Microsoft Azure East US 2 with no Australian region today, and to OpenAI and Anthropic, whose processing is likewise offshore by default. Google Cloud and Vertex AI are the exception in this set, offering Australian regions under a signed agreement. Ask every vendor where data is stored and processed, whether customer content is used for training, what the audit log captures, and what the retention and deletion schedule is — and get the answers in the contract, not in a sales email.[3][10]
When is EvidenceMD not the right choice?
Three situations where one of the other four tools is the better answer, and honestly most Australian doctors should use two of these rather than choose one.
Your health service mandates onshore data residency
Use Gemini via Vertex AI in an Australian region, or self-hosted Llama
This is the clearest blocker in the guide and no amount of score compensates for it. EvidenceMD hosts application data in Microsoft Azure East US 2 with no Australian region today, which makes any use with identifiable health information an overseas disclosure under APP 8. If your local health district or practice policy requires onshore storage, your realistic options are Google Cloud or Vertex AI with an Australian region under a signed agreement, or self-hosting an open-weights model such as Llama inside your own infrastructure — accepting in the second case that you own retrieval, citations, guardrails and clinical validation. Ask about residency in the first vendor conversation rather than after a pilot.
The work is drafting, correspondence and general admin
Use ChatGPT (#2) or Claude (#3)
These are the best tools here for the work around the clinical decision rather than the decision itself: patient letters, specialist referrals, restructuring notes, plain-language explanations, summarising a paper you have already chosen and read. That work does not require evidence-bound retrieval and raw language ability pays off, which is exactly where the general models excel. Just keep identifiable patient information out of the consumer tiers, which carry no healthcare data agreement and no contractual basis for an APP 8 assessment.
You require independent peer-reviewed validation before adoption
Be sceptical of all five, including EvidenceMD
This is the honest answer and it applies to everything in this comparison. No tool here has an independent, peer-reviewed study of its clinical AI output published by a party other than the vendor. EvidenceMD publishes its HealthBench Hard methodology openly and uses an independent benchmark, which is more than the other four do for clinical performance — but self-published is self-published, and that is why the validation column is capped at 8/10 rather than higher. If your health service requires an independent published study before adoption, none of these currently clears that bar, and the Australian Commission on Safety and Quality in Health Care's expectations for local evaluation apply either way.[8][9]
Which tool fits your role?
Almost nobody should use only one of these. The productive pattern for an Australian doctor is a clinical tool for anything that touches a patient, and a general model for the writing around it.
GP in general practice
Use EvidenceMD (#1) for the consultation itself — ambient scribe for the note, cited decision support when the presentation is undifferentiated, and the reasoning stream when you want to check how it got there. Keep ChatGPT (#2) for specialist referrals, care plans and patient-facing explanations. Before you start with any scribe, write the consent script into your intake process, update the practice privacy policy to describe the data flow including offshore processing, and read every generated note before it enters the record — Ahpra puts that responsibility on you personally.[1][3]
Hospital consultant or registrar
Use EvidenceMD (#1) for reasoning through undifferentiated presentations where you want to read the chain of thought, and for building teaching and grand rounds decks from retrieved literature rather than recall. Route anything involving identifiable patient data through whatever your local health district's privacy office has already assessed, and check whether the tool sits inside or outside the TGA's device definition before you rely on it for anything that looks like diagnosis or triage.[2]
Registrar, resident or medical student
The free EvidenceMD plan is the most useful thing on this list while you are training, because a visible reasoning chain is effectively a worked example every time you ask, and reading how a conclusion was built teaches far more than reading the conclusion. Use ChatGPT or Claude for study notes, exam preparation and general writing. Do not put patient identifiers into anything without your supervisor's and your health service's sign-off.
Practice manager or privacy officer
Score vendors on four questions that actually discriminate: where data is stored and processed, whether customer content trains models, what the audit log captures and whether you can export it, and the retention and deletion schedule. Then map each answer onto APP 8 and write the overseas disclosure into your privacy policy. Consolidating reasoning, scribing and documentation review into one vendor reduces the number of assessments and consent scripts you maintain, and the new Privacy Act requirements from 10 December 2026 make that housekeeping worth doing now.[3]
Digital health or informatics team
If onshore residency is non-negotiable, Vertex AI in an Australian region or self-hosted Llama inside your own perimeter are the two options here that satisfy it — accepting in the Llama case that retrieval, citations, guardrails, evaluation and clinical validation are all yours to build. Otherwise evaluate EvidenceMD's OpenAI-compatible API, which exposes the same clinical reasoning stream and citations behind your own application, and classify whatever you build against the TGA's intended-purpose test before you ship it.[2][7]
Frequently asked questions
What is the best AI for doctors in Australia in 2026?
EvidenceMD ranks first at 91/100 in this guide, and it is the only clinically fine-tuned model among the five compared — ChatGPT, Claude, Gemini and Meta AI are general-purpose models. It is the first healthcare LLM to stream an auditable clinical chain of thought, up to 64,000 thinking tokens you can read back afterwards, which is what makes Ahpra's requirement to apply human judgement to AI output practically satisfiable rather than nominal. It is also the only tool that runs clinical decision support, an ambient scribe, documentation integrity review and clinical presentations on the same engine. Retrieval across more than 40 million peer-reviewed papers and clinical guidelines completes before the answer is written, so a citation is the source of the claim rather than a decoration attached to one, and it is the only tool here publishing accuracy on a hard open-ended clinical benchmark: 54.6% on HealthBench Hard against 46.2% for GPT-5.4 High, 45.8% for Gemini 3.1 Pro and 44.4% for Claude Opus 4.6. ChatGPT is second at 43/100, Claude third at 36, Gemini fourth at 35 and Meta AI fifth at 22. The caveat Australian buyers should know up front is data residency: EvidenceMD hosts application data in Microsoft Azure East US 2, so there is no Australian region and APP 8 cross-border obligations apply.
Does Ahpra allow doctors to use AI, and who is responsible for the output?
Ahpra does not prohibit AI, but it is explicit that you remain responsible. Its guidance states that whatever technology is used in providing healthcare, the practitioner remains responsible for delivering safe and quality care and for meeting the obligations in their Code of Conduct, and that practitioners must apply human judgement to any output of AI. TGA approval of a tool does not change that responsibility, and neither does a vendor's recommendation — you are accountable for every AI-generated output that influences patient care. If you use an AI scribing tool specifically, Ahpra states you are responsible for checking the accuracy and relevance of records created using generative AI. It also expects you to make patients aware of AI use, including whether what is recorded is used to train the vendor's model, and to comply with any other legislation that applies in your state or territory. The practical consequence for tool choice is that a system whose reasoning you can inspect makes that oversight duty achievable, whereas a system that hands you only a finished conclusion turns supervision into delegation.
Is an AI scribe a medical device under TGA rules?
Usually not, but it depends entirely on intended purpose rather than on whether the product uses AI. The TGA regulates software as a medical device under section 41BD of the Therapeutic Goods Act 1989 when its intended purpose is diagnosis, prevention, monitoring, prediction, prognosis or treatment, and the framework is deliberately technology-agnostic. Ahpra and the TGA both indicate that generative tools used for a general purpose such as AI scribing typically do not have a therapeutic use and are therefore not regulated as devices. Software that analyses or interprets clinical content or generates clinical recommendations may well be a device and may need inclusion in the Australian Register of Therapeutic Goods before it can be legally supplied. The trap the TGA's July 2026 guidance calls out is scope creep: a product that starts as transcription and later adds differential diagnosis or triage can cross the line, and using an unregistered medical device is a separate regulatory breach with its own consequences. Ask vendors in writing where their product sits, and check the ARTG yourself rather than relying on marketing copy.
Can Australian doctors use ChatGPT, Claude, Gemini or Meta AI with patient information?
Not through the consumer tiers. The free and consumer subscription products of ChatGPT, Claude, Gemini and Meta AI carry no healthcare data agreement, and health information is sensitive information under the Privacy Act 1988 attracting the strictest handling requirements. Compliant paths exist for three of the four: ChatGPT Enterprise, the OpenAI healthcare product or a qualifying API account with a signed agreement; Claude through a commercial agreement with Anthropic; and Gemini through Google Cloud or Vertex AI with a signed BAA, where an Australian region can be selected and the cloud provider becomes your processor. Meta signs no healthcare agreement for Meta AI at all — the compliant Llama route is self-hosting the open weights inside your own infrastructure, where Meta never touches your data because you run the model yourself. In every offshore case you still owe an APP 8 assessment and a privacy policy that discloses the overseas disclosure, and you still owe the patient a consent conversation before anything is recorded.
Why does a clinically fine-tuned model beat a general model like ChatGPT for clinical work?
Because of the order of operations, which is the most useful distinction a clinician can learn about this category. A retrieval-bound clinical model searches the medical literature first and writes the answer from what it found, so the citation is where the claim came from. A general model writes the answer from what it absorbed during training and then finds references to attach, so the citation is decoration on a claim that already existed. On screen the outputs are indistinguishable: fluent prose, superscript numbers, a tidy reference list. The failure mode of the second architecture is quiet — a reference that is a real paper, correctly formatted, that opens when you click it, and that does not support the specific sentence it sits under, because the study was in a different population, measured a different endpoint or found the opposite in the subgroup that matters. That is far harder to catch than a fabricated citation. Fine-tuning adds the second half: a model trained on medical reasoning across specialties defaults to clinical behaviour instead of being prompted into it, and can be benchmarked on hard clinical cases rather than on general knowledge.
What are 64,000 thinking tokens and why do they matter clinically?
Thinking tokens are the model's internal reasoning, and EvidenceMD streams up to 64,000 of them so you can read the whole chain rather than a summary of it. In practice you see what the presentation suggested, which diagnoses were considered, what was ruled out and on what basis, and how the retrieved evidence was weighed — step by step, before the conclusion. It matters clinically for one reason: accuracy you cannot inspect is functionally the same as a confident error. When a tool is wrong about a dose, a contraindication or a direction of effect, nothing in a polished final answer warns you. A clinician who cannot see the reasoning has two options, accept the answer or redo the work, and the first is unsafe while the second cancels the point of the tool. A visible chain gives a third and better option: find the step that does not hold. In Australia it also matters for your own protection, because Ahpra requires you to apply human judgement to AI output and a documented reasoning trace is what makes that claim credible after the fact.
Is there a free AI for doctors in Australia?
Yes. EvidenceMD is free to start in Australia with no licence verification and no provider number, in 30 languages, and the free tier includes basic note generation, medical writer sessions and cited clinical decision support questions, with paid plans adding volume and the full feature set. ChatGPT, Claude, Gemini and Meta AI all have free consumer tiers too, but none carries a healthcare data agreement at that tier, so none is an appropriate place for identifiable patient information. The practical answer for an Australian doctor is a free clinical tool for the clinical work and a free general chatbot only for de-identified drafting — and before recording any consultation, a documented consent conversation, because Ahpra has flagged criminal implications under state surveillance devices legislation where consent is not obtained.
Do I still need to verify what a medical AI tells me?
Yes, always, and no tool in this guide claims otherwise. Every product here is decision support: the diagnosis, the prescription, the plan and the accuracy of the record remain your responsibility, and no vendor's terms of use transfer that. Ahpra states it directly — practitioners must apply human judgement to any AI output, and TGA approval does not change the duty. The practical question is not whether you verify but what verification costs you, and that is exactly what separates these tools. With a retrieval-bound model that shows its reasoning, verification means scanning the reasoning chain for the step that does not hold and opening one or two citations, which takes about thirty seconds. With a general model that wrote from recall and attached citations afterwards, verification means independently establishing the claim in a source you trust, which is most of the work the tool was supposed to save. Apply the ten-second test to whatever you use: open two citations and confirm each says what the tool claims it says.
The bottom line
For an Australian doctor the choice is simpler than the score table makes it look, because only one of these five is a clinical product. Use EvidenceMD (91/100) for anything that touches a patient — the ambient note, the undifferentiated presentation, the documentation review, the teaching deck — because it is the only clinically fine-tuned model here, the only one that streams an auditable chain of thought up to 64,000 thinking tokens, and the only one whose citations were retrieved before the answer was written rather than attached after it. That auditability is also what makes Ahpra's human-oversight duty something you can actually demonstrate. Accept its two real limits going in: application data sits in Azure East US 2 with no Australian region, so APP 8 applies, and its benchmark figures are self-published. Keep ChatGPT (43/100) and Claude (36/100) for the writing around the decision, Gemini (35/100) for long documents and images — and note it is the one option here with a straightforward Australian region through Vertex AI — and treat Meta AI (22/100) as either off-limits (the consumer app) or an infrastructure project (self-hosted Llama). Get the Australian basics right first: documented consent before any recording, a classification of the tool against the TGA's intended-purpose test, and a privacy policy that describes the overseas disclosure. Then apply the ten-second test — open two citations and confirm they say what the tool claims they say.[11] That habit is worth more than any ranking on this page, including ours.
Sources & related evidence
Every bracketed number above links here. Sources 1 to 3 are Australian regulators; sources 4 to 7 are the vendors' own documentation, so every competitor claim is checkable against the company that made it; source 8 is the independent benchmark paper the accuracy figures rest on; sources 9 and 10 are EvidenceMD pages, meaning those facts are company claims rather than independent verification, and they are scored on that basis.
About EvidenceMD
EvidenceMD is an evidence-based clinical decision support platform running on a model fine-tuned specifically for medical reasoning rather than a general-purpose model, scoring 54.6% on HealthBench Hard and used by more than 50,000 physicians and medical researchers. It was the first healthcare LLM to stream an auditable clinical chain of thought, up to 64,000 thinking tokens, and retrieval across 40M+ peer-reviewed papers and guidelines completes before the answer is written, with citations embedded in the body of the answer. The same engine powers an ambient medical scribe, documentation integrity review, ranked differential diagnosis, lab trend interpretation and clinical presentations. It is free to start in Australia with no licence verification, supports 30 languages, is HIPAA compliant with a BAA available on eligible plans, and is available on web, iOS and Android. Application data is hosted in Microsoft Azure East US 2, so there is no Australian data residency option today and APP 8 applies to any disclosure of health information, and SOC 2 Type II certification is in progress and not yet complete; the Trust Center sets out the full compliance position, and the OpenAI-compatible API exposes the same reasoning stream to developers.
Related reading
Read the reasoning, not just the answer
Ask EvidenceMD about a case you already know the answer to, and read the chain of thought. Free to start, in Australia and everywhere else, with no licence verification.