Skip to main content
Cloud Information Model

CIM Formats

The Cloud Information Model (CIM) was designed from the outset as a standards-based, application-agnostic model of business concepts — customers, orders, products, accounts, and the relationships between them. But a conceptual model is only useful if the systems that need it can actually consume it. That is why CIM is published not as a single proprietary artifact but as a family of serializations, each targeting a different class of tooling, runtime, and audience.

This page explains what each CIM format is, what it is for, and how to choose among them. If you are an enterprise data architect, an integration or ETL engineer, an application or platform vendor, or an open-source contributor, the format you reach for first depends on where in the pipeline you sit.

Key Takeaways

  • CIM is distributed in two families: semantic-web formats (JSON-LD, RDF Schema, SHACL, R2RML) and human-readable/relational formats (AML vocabulary, AML dialect, RAML types, JSON Schema, SQL DDL).
  • The conceptual model (concepts.*) describes entities and relationships; the canonical schema (schema.*) describes data shapes and constraints. They are separate artifacts with separate purposes.
  • JSON-LD is the canonical machine-readable form; AML is the human-readable form of the same content; SQL DDL and JSON Schema are the forms most application and ETL teams consume directly.
  • R2RML is the bridge: it maps a relational schema to an RDF graph, which is how you connect existing SQL databases to the semantic layer.
  • Choosing a format is a question of consumer, not preference — pick the one your target toolchain natively ingests, and use the others as cross-checks.

Why CIM Ships in Multiple Formats

Most data models are published in exactly one form — usually an ER diagram, a spreadsheet, or a vendor-specific metadata file. That works until you need to share the model across organizations that use different stacks.

A retail platform might run PostgreSQL and dbt; a partner might run a graph database and a triple store; a SaaS vendor might expose JSON APIs and validate payloads with JSON Schema. If the shared model only exists in one of those dialects, everyone else has to translate it — and translations drift.

CIM’s multi-format strategy is a deliberate answer to that problem. The model is authored once and then translated into formats that map cleanly onto recognized standards, so that each consumer adopts CIM using tools it already has.

This is the same philosophy behind standards bodies such as the World Wide Web Consortium (W3C), which publishes specifications like RDF, SHACL, and R2RML that CIM reuses rather than reinvents. It also aligns with the broader interoperability mission of the Linux Foundation, under which the CIM project operates.

The practical benefit is twofold: businesses with varying technologies can adopt CIM without a rip-and-replace, and contributors can extend the model in whichever format matches their expertise, knowing the other serializations can be regenerated.

The Conceptual Model vs. the Canonical Schema

Before comparing file formats, it helps to separate two layers that CIM keeps distinct — and that newcomers frequently conflate.

  • The conceptual model answers what exists and how it relates. It defines entities (Customer, Order, Product), their attributes, and the relationships between them. It is intentionally close to business vocabulary and deliberately light on physical detail.
  • The canonical schema answers what a valid instance looks like. It adds data shapes and constraints — cardinality, types, required fields, value ranges — that a system can validate against.

In the CIM distribution, these map to two filename stems: concepts.* for the conceptual layer and schema.* for the canonical layer. Keeping them separate means a business analyst can read the conceptual model without wading through constraint syntax, while an engineer can validate payloads against the schema without needing the full conceptual narrative.

The Semantic-Web Formats

These formats express CIM as an RDF-based graph. They are the right choice when your consumers include triple stores, knowledge graphs, ontology tooling, or any system that reasons over linked data.

JSON-LD — concepts.json and schema.json

JSON-LD is JSON with a linked-data context, which makes it the pragmatic bridge between ordinary web APIs and the semantic web. CIM publishes two JSON-LD artifacts:

  • concepts.json — the conceptual description of entities and relationships, expressed as RDF Schema.
  • schema.json — the canonical data shapes and additional constraints, expressed in SHACL.

Because it is valid JSON, concepts.json and schema.json can be loaded by ordinary JSON tooling, but because they carry an @context, they also expand into full RDF triples. That dual nature is why JSON-LD is often the best default for teams that want semantic fidelity without adopting a specialized RDF stack on day one.

RDF Schema — schema.json

RDF Schema (RDFS) provides the vocabulary for describing classes and properties — the rdfs:Class, rdfs:subClassOf, and rdfs:domain/rdfs:range constructs that let a machine understand that an Order is a business document and that its customer property points at a Customer. CIM uses RDFS to give the conceptual model formal semantics, so that subclass hierarchies and property domains are machine-interpretable rather than merely documented.

SHACL — schema.json

The Shapes Constraint Language (SHACL) is a W3C standard for validating RDF graphs against a set of conditions called shapes. Where RDFS says what a class is, SHACL says what a valid instance must satisfy — required properties, allowed value types, cardinality limits. CIM’s canonical data shapes are expressed in SHACL, which means any SHACL processor can validate CIM-conformant data without custom code.

R2RML — schema.rdml

R2RML is the W3C standard for mapping a relational database schema to an RDF graph. This is the format that matters most to integration and ETL engineers, because it is the mechanism by which an existing SQL database — with its tables, columns, and foreign keys — is exposed as CIM-conformant linked data.

Rather than re-modeling your operational database by hand, you write (or generate) an R2RML mapping that declares how each table and column corresponds to CIM entities and properties. The result is a virtual RDF graph over your existing relational data.

The Human-Readable and Relational Formats

Not every consumer wants RDF. Application developers, data modelers, and DBAs often want something they can read in a text editor or load straight into a database. CIM serves them with the AML, RAML, JSON Schema, and SQL DDL serializations.

AML — concepts.yaml, schema.yaml, schema.raml

AML, the AnyLogic Modeling Language lineage used here as a modeling dialect, is the human-readable expression of CIM. CIM publishes three AML artifacts:

  • concepts.yaml — the AML vocabulary, a human-readable version of the conceptual model.
  • schema.yaml — the AML dialect, a human-readable version of the canonical data shapes.
  • schema.raml — the RAML data types rendering of the canonical shapes.

The distinction between vocabulary and dialect is worth internalizing: the vocabulary defines the terms (the nouns and verbs of the model), while the dialect defines how those terms are combined into valid structures. If you are reviewing CIM for the first time, concepts.yaml is usually the most approachable entry point.

JSON Schema — schema.json

JSON Schema is the de facto standard for validating JSON documents, supported natively or via libraries in virtually every modern language. CIM’s JSON Schema artifact expresses the canonical data shapes as JSON Schema, which makes it directly usable in API gateways, message brokers, and CI pipelines that already validate JSON payloads. If your integration surface is REST or event-driven JSON, this is often the format you want.

SQL DDL — schema.sql

SQL DDL is the set of CREATE TABLE, CREATE VIEW, and constraint statements that materialize the canonical shapes in a relational database. CIM targets SQL 2008 syntax, which keeps the DDL portable across the major relational engines. This is the format DBAs and ETL engineers reach for when they want to stand up a physical schema that conforms to CIM — for example, a staging or integration database that mirrors the canonical model.

Choosing a Format: A Practical Guide

There is no single “correct” format. The right choice is determined by who or what consumes the model next. Use the table below as a decision aid.

If your consumer is…Start with…Because…
A business analyst or data modeler reviewing the modelAML vocabulary (concepts.yaml)Human-readable, business-vocabulary-first
A triple store, knowledge graph, or ontology toolJSON-LD (concepts.json, schema.json)Native RDF with a JSON on-ramp
A SHACL validator or semantic data-quality pipelineSHACL (schema.json)Standard constraint validation over RDF
An existing relational database you want to expose as linked dataR2RML (schema.rdml)Maps tables/columns to CIM entities without re-modeling
A REST or event-driven JSON APIJSON Schema (schema.json)Directly validates JSON payloads
A relational database you want to conform to CIMSQL DDL (schema.sql)Portable SQL 2008 DDL
A RAML-described APIRAML types (schema.raml)Native to RAML toolchains

A few practical caveats:

  • Do not treat the formats as independent models. They are serializations of the same underlying CIM. If you find a discrepancy between, say, schema.json (JSON Schema) and schema.json (SHACL), that is a bug or a version skew, not a design choice — report it.
  • Watch the filename collisions. Several formats share the stem schema with different extensions (schema.json, schema.yaml, schema.raml, schema.sql, schema.rdml). When downloading the full distribution, keep the format directories separate so you do not overwrite one serialization with another.
  • Match the format to the validation stage. Use the conceptual formats for design-time review and the canonical formats for runtime validation. Validating against the conceptual model is not meaningful — it lacks the constraints.
  • Prefer generated over hand-edited. If you extend CIM, extend the source and regenerate the other serializations rather than editing each format by hand, or the family will drift out of sync.

Downloading the Full CIM Distribution

CIM is distributed as a complete definition in each available format, so you can pull down the entire model in the serialization you need rather than assembling it piece by piece. The published download options are:

  • AML (vocabulary) — the human-readable conceptual model.
  • AML (dialect) — the human-readable canonical shapes.
  • JSON-LD (vocabulary & schema) — the machine-readable semantic model.
  • R2RML — the relational-to-RDF mapping.
  • RAML Types — the canonical shapes as RAML data types.
  • SQL DDL — the canonical shapes as portable SQL.

Each download contains the full CIM definition in that format, which means you can adopt CIM incrementally: start with the format your current toolchain supports, and add others as your interoperability needs grow.

Contributing Across Formats

Because CIM is an open project, contributions are welcome — and the multi-format structure shapes how contributions work. Contributors typically fall into two groups:

  • Model contributors propose new entities, relationships, or constraints. These changes are authored once and then propagated to the other serializations.
  • Format contributors improve the fidelity or tooling of a specific serialization — for example, refining the R2RML mappings or the SQL DDL portability.

If you are contributing, the practical rule is to understand which layer you are changing (conceptual vs. canonical) and which formats must be regenerated as a result. The project’s GitHub repositories and contributor web form are the entry points for getting involved.

Frequently Asked Questions

What is the difference between concepts.json and schema.json in CIM?

concepts.json is the conceptual model — the entities and relationships in CIM, expressed as JSON-LD with RDF Schema semantics. schema.json is the canonical schema — the data shapes and additional constraints, expressed as JSON-LD with SHACL semantics. In short, concepts describes what exists; schema describes what a valid instance must look like.

Why does CIM publish the same model in so many formats?

Because different consumers use different technologies. A triple store needs RDF; a JSON API needs JSON Schema; a DBA needs SQL DDL; a business analyst needs something human-readable. Publishing CIM in multiple standard formats lets each of those audiences adopt the model with tools they already have, rather than forcing a single stack on everyone.

What is R2RML used for in CIM?

R2RML is the W3C standard for mapping a relational database schema to an RDF graph. In CIM it is the bridge that exposes existing SQL databases as CIM-conformant linked data, so you can connect operational relational systems to the semantic layer without re-modeling them by hand.

Is AML the same as the JSON-LD formats?

No. AML is the human-readable expression of CIM — the vocabulary (concepts.yaml) and dialect (schema.yaml, schema.raml). JSON-LD is the machine-readable, RDF-based expression. They describe the same model but target different audiences and toolchains.

Which CIM format should I start with?

It depends on your consumer. If you are reviewing the model, start with the AML vocabulary. If you are building a JSON API, start with JSON Schema. If you are connecting a relational database, start with R2RML or SQL DDL. If you are working with a knowledge graph, start with JSON-LD and SHACL.

Can I edit one CIM format without updating the others?

You can, but you should not. The formats are serializations of one underlying model, so editing a single format by hand causes the family to drift out of sync. Extend the source model and regenerate the other serializations instead.

Further Reading

Frequently asked questions

What is the difference between `concepts.json` and `schema.json` in CIM?

concepts.json is the conceptual model — the entities and relationships in CIM, expressed as JSON-LD with RDF Schema semantics. schema.json is the canonical schema — the data shapes and additional constraints, expressed as JSON-LD with SHACL semantics. In short, concepts describes what exists; schema describes what a valid instance must look like.

Why does CIM publish the same model in so many formats?

Because different consumers use different technologies. A triple store needs RDF; a JSON API needs JSON Schema; a DBA needs SQL DDL; a business analyst needs something human-readable. Publishing CIM in multiple standard formats lets each of those audiences adopt the model with tools they already have, rather than forcing a single stack on everyone.

What is R2RML used for in CIM?

R2RML is the W3C standard for mapping a relational database schema to an RDF graph. In CIM it is the bridge that exposes existing SQL databases as CIM-conformant linked data, so you can connect operational relational systems to the semantic layer without re-modeling them by hand.

Is AML the same as the JSON-LD formats?

No. AML is the human-readable expression of CIM — the vocabulary (concepts.yaml) and dialect (schema.yaml, schema.raml). JSON-LD is the machine-readable, RDF-based expression. They describe the same model but target different audiences and toolchains.

Which CIM format should I start with?

It depends on your consumer. If you are reviewing the model, start with the AML vocabulary. If you are building a JSON API, start with JSON Schema. If you are connecting a relational database, start with R2RML or SQL DDL. If you are working with a knowledge graph, start with JSON-LD and SHACL.

Can I edit one CIM format without updating the others?

You can, but you should not. The formats are serializations of one underlying model, so editing a single format by hand causes the family to drift out of sync. Extend the source model and regenerate the other serializations instead. Further Reading - [World Wide Web Consortium (W3C)](https://www.w3.org/) — the standards body behind RDF, RDF Schema, SHACL, and R2RML, the specifications CIM builds on. - [Shapes Constraint Language (SHACL)](https://en.wikipedia.org/wiki/SHACL) — background on the constraint language used for CIM's canonical data shapes. - [Linux Foundation](https://en.wikipedia.o


More on Open source enterprise data interoperability standard / cloud data modeling

Browse our latest guides and reviews.

Read more →