Large Objects |
![]() |
CDO large objects are immutable, identifiable values for content that should not be serialized as an ordinary
in-line model attribute. A CDOBlob represents binary content and a CDOClob represents character
content. A model revision stores the large object's identity and size; the content is handled by the CDO large-object
store and is read through a stream or reader when needed.
This chapter explains the client-facing model, creation and consumption APIs, transaction behavior, cache-backed loading, model representation, and resource ownership. It is about CDO large objects rather than JDBC BLOB/CLOB objects. General session, view, and transaction lifecycle remains in Working with Sessions, Working with Views, and Working with Transactions.
Table of Contents
CDO's public large-object types are CDOBlob and CDOClob. They are the Java values for the
Blob and Clob EDataTypes in the CDO types package. A
generated model feature using one of those data types has a Java type of CDOBlob or CDOClob; the
value is not a byte[], a String, or a mutable file handle.
A byte array or string attribute is part of the ordinary revision value and is normally materialized with that revision. A large-object attribute is represented by a content identifier and size, while the content is stored and transferred separately. This separation avoids putting the complete payload into every revision serialization.
Both large-object classes are immutable. CDOLobInfo.getID() identifies the content with a digest and
CDOLobInfo.getSize() reports its size. The size is bytes for a blob and characters for a clob. Treat the
returned ID array as read-only. The public equality and hash-code contract is inherited from CDOLobInfo:
it compares the ID bytes and size, rather than reading and comparing the content.
Create a blob from an InputStream when the source is naturally streamed, or from a byte[] when the
content is already in memory. A hexadecimal string is also accepted by the low-level CDOBlob constructor,
but it represents encoded binary data, not ordinary text. Prefer the session factories
CDOSession.newBlob(InputStream) and CDOSession.newBlob(byte[]) because they use the session's
configured large-object cache.
Construction consumes the supplied stream immediately to compute the identity and populate the client large-object store. The constructor does not retain the stream for a later commit. The caller owns the supplied stream and should close it, normally with try-with-resources.
Create a clob from a Reader for streamed character input or from a String for already-materialized
text. Use CDOSession.newClob(Reader) and CDOSession.newClob(String) so that the session's configured
large-object store is used. Character decoding belongs at the boundary: if the source is bytes, choose its charset
when creating the reader. A clob's size is counted in characters, not source bytes.
As with blobs, construction consumes the supplied reader immediately and does not retain it. The caller owns and closes the supplied reader.
A large object is assigned to a model feature in the same way as any other EDataType value. The assignment changes
the transaction's local model and makes the transaction dirty; the value becomes visible to other views only after
a successful commit. Replacing a blob or clob is an ordinary feature replacement
from the transaction's perspective.
The payload has already been consumed into the client large-object store when the value is constructed. During commit, CDO sends the content needed by the repository for large-object identities that are not already known there, together with the revision change. Consequently, an I/O failure can occur during creation as well as during commit. Rolling back removes the feature change from the transaction; it does not imply that every temporary local cache file is immediately deleted.
Content identity allows a repository to recognize an already-known large object and avoid storing or transferring a duplicate payload. This is an identity optimization, not a reason to mutate a large object: large-object values are immutable, so replacing content means assigning a new value.
See Working with Transactions for dirty state, commit, rollback, and conflict handling.
CDOBlob.getContents() returns an InputStream; CDOClob.getContents() returns a
Reader. Both can be requested repeatedly. The returned resource belongs to the caller and must be closed.
The convenience methods CDOBlob.copyTo(OutputStream), CDOClob.copyTo(Writer),
CDOBlob.getBytes(), and CDOClob.getString() are useful when their materialization cost is
acceptable. For large payloads, prefer copying from the stream or reader to the final destination.
A model object or revision can be loaded without reading the full large-object content. On the first content access,
a session-backed store can load missing content from the repository into its configured cache and then return a
stream or reader. Subsequent reads can use that cache. The exact cache behavior is determined by the configured
session LOB cache; do not assume that a custom store is disk-backed or
reusable after its owner has been closed. Keep the session available while a missing remote payload is being read.
The revision value contains large-object identity and size, while the bytes or characters are obtained through the
large-object store. The default client store is file-backed, and a session exposes its configurable
LOB cache through session options. If a requested value is absent from the
cache, the session's store can fetch it from the repository before returning the content stream or reader.
The cache is keyed by the large-object identity and is separate from the ordinary object/revision cache. It can
therefore reduce repeated network transfers without making a blob or clob mutable. Cache location, retention, and
eviction are properties of the selected CDOLobStore; applications should
choose a suitable store when local disk or memory usage matters.
Loading a revision is not the same as loading every large-object payload referenced by that revision. Avoid calling
CDOBlob.getBytes() or CDOClob.getString() merely to inspect metadata; use CDOLobInfo.getID()
and CDOLobInfo.getSize() instead.
Large-object support is a store capability. CDO's public client API does not expose a universal boolean capability
query on CDOCommonRepository; applications should therefore treat large-object
support as a repository/store deployment requirement and verify it for the repositories they target. A store that
does not support large objects cannot provide the normal commit and load behavior.
The repository also exposes its configured LOB digest algorithm. Applications normally do not need to depend on that value, but it explains why the ID is a
content-derived identity rather than an arbitrary object ID. Server/store cleanup of unreferenced large objects is a
store concern and is not part of ordinary client transaction lifecycle.
Declare a persistent EAttribute with the CDO Blob or
Clob EDataType when a generated model feature should hold a CDO large object. The
generated accessor then uses CDOBlob or CDOClob. The built-in CDOBinaryResource and
CDOTextResource are examples of model objects whose persistent contents feature has exactly this shape.
These values are attributes, not EObjects: do not put them in containment, and do not model their payload as an ordinary byte-array or string feature if the application needs CDO large-object storage and streaming. Lists of blob or clob values are supported by the same value semantics; each entry is still immutable and identified by its own content ID. General model preparation is covered by the Preparing Models material rather than repeated here.
A large object's ID is a digest of its content, with the repository's configured digest
algorithm used consistently by the client and store. The ID also distinguishes binary and character content in the
standard store. CDOLobInfo.equals(Object) compares the ID bytes and size; it does not stream the payload.
This makes equality inexpensive and lets CDO recognize an already-known content value during commit.
Equality is not Java object identity: two separately constructed values with the same content can compare equal. Conversely, do not compare a blob and clob merely because their visible text happens to match; they are different large-object kinds and have different content semantics. Treat the ID array returned by the API as immutable even though the array itself is exposed.
Close input streams and readers supplied to blob/clob constructors or session factories after construction has
completed. Those factories consume the source synchronously but do not make the source application's lifecycle their
responsibility. Always close streams and readers returned by CDOBlob.getContents() and
CDOClob.getContents(), even when copying fails.
Creation can fail with IOException while reading the source or writing the client store. Reading can fail
with IOException because cached content is missing, a repository transfer fails, or the returned stream or
reader cannot be consumed. A closed or disconnected session is especially relevant when the needed content is not
already in the local cache. Handle these failures as I/O and repository failures; do not assume that model-object
loading has already validated the payload.
Keep only metadata when possible. Holding a byte[] or String obtained from CDOBlob.getBytes()
or CDOClob.getString() defeats the memory benefit of the streaming API.
Use a CDO large-object feature when content is large enough that embedding it in ordinary revision values would
create undesirable materialization, serialization, or memory costs. Construct from streams/readers for file or
network sources, and consume through streams/readers for file or network destinations. Use byte arrays, strings,
CDOBlob.getBytes(), or CDOClob.getString() only when the complete content is intentionally needed.
Remember that construction performs a complete local read to calculate identity and cache the content; streaming creation avoids an extra application buffer but does not make digesting free. Commit may transfer content that the repository does not already know, and the first read of a missing value may transfer it back. Reuse the immutable large-object value or its cache when appropriate, close resources promptly, and configure a suitable session LOB cache for the application's disk and memory constraints.
Choose a blob for binary data and a clob for character data. Decode bytes with an explicit charset before creating a clob, and do not use a hexadecimal blob constructor as a text encoding shortcut.