Skip to main content
Chemia Discovery

Reading results

Some fields mean what they say and some do not. We sort each reading note into one of four classes: Obvious, Convention, Trap or Gap. As of 2026-10-04 we publish 64 notes: 35 Conventions, 21 Traps and 8 Gaps.

At a glance
ClassesObvious, Convention, Trap, Gap
Notes published64
Conventions35
Traps21
Gaps8
Tools with notes28 of 28
Data as of2026-10-04
Server release1.5.0, read from the live server on 2026-10-04

What are the four reading classes?

The four reading classes (as of 2026-10-04)
ClassWhat it meansNotes published
ObviousA field that means what its name says. It needs no note, so we publish none.none, by definition
ConventionA field or behaviour whose meaning is not obvious.35
TrapSomething that looks right and is not.21
GapSomething we do not hold or do not check.8

Which fields are Obvious?

A field is Obvious when its name tells you what it holds and nothing about it is easy to misread. Obvious fields carry no note, so the snapshot lists none. The worked examples will tag every field they show with its class, Obvious included. Examples says where they stand.

What are the Conventions?

A Convention is a field or behaviour whose meaning is not obvious. These 35 notes are published as of 2026-10-04, by tool.

  • benchmark_against: The five candidates are the first five rows of the set in its current order. The incumbent is read from the same dataset the set came from. Normalized values use each property's realistic range, oriented the way the set was ranked.
  • classify_material: The class follows the band gap. A gap of zero is metal, and so is a metal flag when no positive gap is stored. A gap above zero and below 0.1 eV is semi-metal; from there up to, but not including, 3 eV is semiconductor; 3 eV and above is insulator. A material with no band gap value is "unknown". Superconductors are not classified.
  • classify_material: A formula with several structures is classified across all of them. When every entry that has data agrees, the status is "found" with classification_basis all_polymorphs_agree and no single Chemia ID. When they disagree the status is "ambiguous" with class_breakdown and the ground_state's class; pass a Chemia ID to classify one structure.
  • compute_pourbaix_stability: The value is computed, not measured, so get_material_property cannot return it. The conditions (pH window, potential, temperature) and the screening verdict are in note. caveat always travels with it: this is an equilibrium tendency, not a rate, a lifetime or a durability claim. attribution credits the Materials Project data.
  • compute_pourbaix_stability: Each call is charged against the daily allowance and needs the materials_project_data grant, because it calls the live Materials Project API on the platform's shared allowance. A search that asks for water inertness is gated by the same grant.
  • count_materials: count is None when the count did not run, and exactly one of not_countable, clarification_needed or property_offer says why. A count of 0 is a real answer; None is not one.
  • count_materials: Every call is charged against the daily allowance, including a call with empty criteria. The charge is taken before the work runs and is refunded only when the reply is a clarification question.
  • export_result_set: The file is not returned inline. On the hosted server you get a download_url that works in a browser without signing in for 15 minutes, and a resource_uri that returns the same content through the MCP resources/read request. row_count is the number of stored rows.
  • find_material_entries: Rows come most stable first (lowest energy above hull; entries with no stability energy last). ground_state is true for the formula's entry on the hull. When the dataset holds no hull entry for the formula, lowest_energy_in_source marks its most stable entry instead, which is not the ground state. Neither flag is set for a Chemia ID lookup or a partial match.
  • find_property_proxies: verdict is backfillable_now, rankable_with_scatter, elimination_only, irreducible or held_directly. Proxies rank and screen; they never qualify. Only a backfillable route estimates a value, and its error is the input error. A rankable route gives an ordering with stated scatter, an elimination route can only rule materials out, and neither replaces a measurement. held_directly means the corpus stores the property under its own name: read it instead.
  • get_all_material_properties: Each property is chosen by the same rules as get_material_property (experimental, then DFT, then ML; oldest row on a tie), with the same caveat and suspect fields. count is the number of properties returned.
  • get_material_property: When a property has several stored values, one is returned, chosen by the kind of evidence: experimental first, then DFT, then ML. Any other source type, such as literature, is used only when nothing ranked exists, and a tie goes to the oldest stored row. evidence_source says which kind you got; you do not get the range of values.
  • get_material_property: caveat appears when a DFT value comes from a method with a known systematic error, for example a DFT band gap that typically underestimates the measured one. The number is returned as stored, not corrected. A value carrying suspect is physically impossible for that property and is a data error, not a measurement.
  • get_result_set: total_survivors is the true match count of the original search. retrievable is how many matches were stored, up to 5000, and is the number you can page through. truncated_at_persist is true when the search matched more than were stored; no offset reaches those rows.
  • graph_neighbors: Rows are ordered by distance, then by the weight of the direct edge, which is how often the literature links the two. Only neighbours one step away carry a relation and a weight; farther rings have an empty relation and a weight of 0.
  • graph_shortest_path: The path follows edges in their stored direction, from source to target. A path that exists only the other way is not found, so try swapping the two.
  • graph_top_connected: Entities are ranked by their number of distinct neighbours, counted when the graph was loaded. Each row's material is an entity name, which can be an element, an application or a property as well as a material.
  • list_applications: The ranking is a literature co-occurrence count. It says how often a material and an application appear together, not how well the material performs in it.
  • list_material_sources: restricted means reading the dataset needs the professional_database grant. is_default marks the dataset a search would use for you when neither a source_id nor the query text names one, so it can differ between accounts. specialized marks a slice of the full catalogue chosen by application.
  • ontology_expand: derived_property_filters are operator and value pairs read from graph text, with the property word as MatKG wrote it. Treat them as hints for a query; they are not checked against Chemia's property keys or against the corpus.
  • ontology_substitutes: cooccurrence_score is relative to the best candidate for this incumbent, so the top candidate always scores 1. Candidates below a fixed cutoff are dropped. It measures shared literature applications, not similarity of properties: check a candidate with get_material_property or search_materials before recommending it.
  • plot_distribution: bins holds 20 equal-width bins between the 1st and 99th percentile, plus one wider bin at each end that holds every value beyond them. Draw the end bins as open-ended: a few corrupt extremes would otherwise flatten the chart.
  • plot_distribution: stats counts stored values, not materials. A material with values from several kinds of evidence counts once per value, so value_count can exceed material_count, and the percentiles are over values.
  • plot_knowledge_graph: The tool reads no data. It shapes the entities you pass in into nodes and edges, and checks nothing against MatKG or the corpus.
  • plot_results: A search's own result set is unscored, so a spider chart on it fails with no axes. Run score_results first and chart the set it returns. The spider axes are the properties the set was scored on.
  • plot_results: In an Ashby plot a point is on the frontier when no other point is at least as good on both axes and better on one, using each axis's direction. Only points with a value on both axes are plotted. Frontier points come first, and points_total says how many exist when fewer are returned.
  • property_stats: The statistics describe every stored value, from every kind of evidence, while search_materials and count_materials test one value per material (the best-ranked source). A material with a DFT value and an ML value counts twice here, so value_count can exceed material_count and the percentiles are over values, not materials.
  • refine_result_set: Refining works on the stored matches of the named set, which is at most 5000 rows, and never queries the database again. Use show_eliminated on the new set to see what fell out and why.
  • rerank_result_set: There are two different operations. sort_by is a true single-column sort. primary_constraint and weights re-weight a combined score, so the result is not sorted by any one column when the set has other axes. Exactly one of the three must be given, and an empty weights object counts as not given.
  • rerank_result_set: Materials with no value for the sort property are listed last in both directions, because unknown is not the lowest value. The reply's note says how many were left unmeasured.
  • resolve_material_name: The tool answers with one material by design. A formula that fits several structures returns status "ambiguous" with up to 8 candidates, the full total and the ground_state. The candidates are in Chemia ID order, not stability order, so the first is not the most stable; find_material_entries lists every entry.
  • score_results: The score is a weighted mean of per-property scores between 0 and 1. For a limit such as "above X" a property scores 0.5 at X and rises with the margin past X, counted as a fraction of the property's realistic range; "below X" mirrors it, and inside a "between" range the score is 1. A property added through weights has no limit, so it is scored from 0 to 1 across its realistic range in the direction the vocabulary prefers. The scores are not returned: the row order is the ranking.
  • search_materials: There are three different counts. total_survivors is the exact number of matches. results holds at most 50 of them, and truncated is true when it holds fewer than the total. Up to 5000 matches are stored for get_result_set; rows beyond that were counted but never stored. A query judged to match that many or more, or a scan that runs out of time, is not answered with rows: the reply is clarification_needed with a size estimate and ways to narrow it, and the call is refunded.
  • show_eliminated: total is how many materials were eliminated; stored is how many the set kept to show, which for a search is the closest 50 per constraint. total, not stored, answers "how many failed".
  • show_eliminated: Candidates are sorted by violation_amount, the shortfall in the property's unit to four significant figures. A candidate removed for lacking a value has no violation amount and is listed after every measured near miss.

What are the Traps?

A Trap looks right and is not. These 21 notes are published as of 2026-10-04, by tool.

  • benchmark_against: Where the incumbent has no data, its value is raw null with normalized 0.0. That is "no data", not a score of zero; read raw.
  • classify_material: The band gap used is the largest value in the best evidence tier, not the oldest value that get_material_property returns, so the two can differ for a material with several stored values. Separately, a short fixed list of well-known semiconductors whose DFT gap collapses to zero is reported as semiconductor, with a caveat; the raw DFT evidence stays visible in the row.
  • export_result_set: The link is the credential. Anyone who holds the URL can download the file until it expires, so do not paste it where others can read it.
  • find_material_entries: A match_type of partial is a substring fallback. Its rows only contain the text you typed, so they can be unrelated compounds.
  • find_property_proxies: material_id is compared with the internal id column. The Chemia ID shown to users is not accepted there; use the material_id from a result row.
  • get_all_material_properties: "Most stable" is the lowest energy above hull among the formula's entries in the dataset you query. A smaller dataset can lack the true ground state, in which case the answer is the best entry it holds.
  • get_material_property: A property name that is not the stored key is not an error. It reads as status not_found, "no measured value", which looks like missing data. Use the canonical key; property_stats and plot_distribution return did_you_mean for names they do not know.
  • get_result_set: Owning a result set is not enough. A set that came from a restricted dataset stops being readable if your account no longer holds the professional_database grant, even though the rows were yours when the search ran.
  • graph_neighbors: An empty list does not say why. It can mean the name is not in the graph or that the entity has no neighbours. If the graph is not loaded the reply carries a note saying so. The graph is shared by every dataset: it is not scoped to trial or production.
  • graph_shortest_path: A length of 0 with no steps does not say why. It covers an unknown name, no path within 6 edges, a search that grew past its budget, and a source and target that resolve to the same entity.
  • ontology_expand: confidence is the average confidence of the derived property filters, not a probability that the candidates are right. It is 0.0 whenever no filter was derived, even if candidate materials were found.
  • ontology_substitutes: An application filter that matches nothing is silently ignored, and the substitutes are found across all the incumbent's applications. Check shared_applications on each candidate.
  • ontology_substitutes: db_match is the first dataset entry whose formula equals the candidate's name exactly, not the most stable structure, and it carries no Chemia ID. It is null when the name is not a formula in the dataset.
  • plot_distribution: min_value and max_value are exclusive and filter stored values. A value exactly equal to a bound is left out, and a material is kept only for the values inside the range, not for having any value there.
  • plot_knowledge_graph: An edge drawn at 1.0 because its weight was missing or not a number is a placeholder, not a weight. Do not read it as data.
  • plot_results: A missing value is reported as raw null with normalized 0.0. Read raw to tell "no data" from "the worst value"; the normalized number alone cannot.
  • property_stats: The minimum and maximum can be physically impossible values, because the corpus holds some data errors. outlier_note says when they look unreliable. Judge a threshold by the percentiles, not by the extremes.
  • refine_result_set: Thresholds are compared with the stored value in the property's canonical unit. The unit you send is only a label, so a value given in another unit filters wrongly and raises no error.
  • resolve_material_name: A common name is pinned to one specific structure. When that structure is not in the dataset you queried, the lowest-energy entry of the formula in that dataset stands in, so the Chemia ID returned can be a different structure than the name usually means.
  • score_results: A material with no value for an axis is skipped on that axis rather than marked down, so a material with fewer measured properties can outrank one with data on every axis.
  • search_materials: Check source in every response. A source_id that is misspelt or not in the catalogue raises no error: it resolves to the registry default, the small trial subset, and requested_source then names that default too, so only a comparison with the id you sent shows the swap. A restricted dataset that was inferred from the query text, or taken as the server default, is narrowed to the trial subset without an error when your account lacks the professional_database grant; there requested_source differs from source.

What are the Gaps?

A Gap is something we do not hold or do not check. These 8 notes are published as of 2026-10-04, by tool.

  • compute_pourbaix_stability: Status not_computable means we hold nothing to compute from: there is no Materials Project solid entry at that composition, the formula has no element other than hydrogen and oxygen, or the name did not resolve to one material. No value is invented. A status such as mp_auth_error instead means the calculation could not run, which is a server problem and not a fact about the material.
  • count_materials: Criteria that SQL cannot evaluate (cost, material class, form factor, operating conditions, and properties computed at query time) are not countable. They come back as not_countable; search_materials can evaluate them.
  • export_result_set: The file holds the stored rows only, up to 5000. truncated_at_persist is true when the search matched more, and the file cannot contain them.
  • find_property_proxies: A route is usable only when every hop's input property has at least one stored value in scope; we set no higher bar for "dense". Routes through a property we hold nothing for are returned under unusable_routes, not dropped, and a property with no usable route is irreducible: it needs literature extraction or direct measurement.
  • list_applications: Only materials with application coverage in MatKG return rows. An empty list means no coverage or an unknown name, and cannot tell the two apart.
  • list_capabilities: Some entries describe capabilities of the wider platform that this server cannot call. An entry with an mcp_tool_name can be called here; an entry without one is described but not callable from this server.
  • ontology_expand: When the concept is not in the graph, or the graph is not loaded, every list is empty and confidence is 0.0.
  • search_materials: A search can answer only part of a question without saying so. A condition the parser cannot map to a property it knows (for example coercivity) is dropped: it is not in the parsed constraints, it has no constraint_ledger entry, and the matches reflect only the other conditions. Its only trace is a sentence in constraints.open_questions, and clarification_needed and notice stay empty. Compare the parsed constraints with what you asked. A ledger entry with state not_supported is the intended signal for a condition that is parsed but cannot be checked; extending it to dropped conditions is a planned fix, not current behaviour. Today it appears only for a parsed constraint whose property is outside the vocabulary or whose kind this server does not enforce (such as a cost ceiling or a material class); it never carries a count. Entries with state no_data count materials left out only because they have no measurement of that property.

What should I do with these notes?

Before you rely on a tool's output, read its notes. A tool with a full page lists them there; the others are listed on this page. Read the Traps first, because those are the cases where a plausible answer is wrong. Treat a Gap as a limit on what we hold or check, and read its note before you draw a conclusion from an empty or zero result. The Tool reference links every tool page.