Skip to content

physicell_reader

tissue_simulator.physicell_reader

PhysiCell snapshot reader and spatial statistics computation.

This module parses PhysiCell agent-based model output (CSV snapshots and .mat snapshots) and recomputes the same node/edge/neighbor schema used internally by GraphColorizer (node_counts, edge_counts, neighbor_dist). It is intended for comparing ABM-evolved tissue states against initial-condition targets generated by this package.

The :func:stats_to_target_statistics adapter converts the output of :meth:PhysiCellReader.compute_spatial_stats into the :class:~tissue_simulator.replicate_generator.TargetStatistics dataclass expected by :class:~tissue_simulator.replicate_generator.ReplicateGenerator.

CREDITS

The PhysiCell .mat snapshot layout and the cell-definition XML schema used by :meth:PhysiCellReader.load_snapshot_mat are documented and reference-implemented by the PhysiCell-Tools python-loader project (pyMCDS / pcdl), distributed under the BSD-3-Clause license:

https://github.com/PhysiCell-Tools/python-loader

When the optional pyMCDS / pcdl package is installed, the reader delegates to it for full-fidelity parsing. Otherwise it falls back to a direct scipy.io.loadmat parser based on the documented PhysiCell 1.10+ default cells-matrix column layout (see load_snapshot_mat). Neither pyMCDS nor pcdl is a required dependency of tissue_simulator; both are optional and only used when present.

PhysiCellReader

PhysiCellReader(default_radius: float = _DEFAULT_RADIUS)

Reader for PhysiCell snapshot output.

Parses CSV and (optionally) .mat snapshots emitted by PhysiCell and recomputes the spatial statistics consumed by tissue_simulator (node_counts, edge_counts, neighbor_dist).

Initialize the reader.

Parameters

default_radius : float Radius to assign when neither radius nor volume is present in the snapshot. Defaults to a PhysiCell-like value (~8.41 micrometers, the radius of a 2494 cubic micrometer cell).

Source code in tissue_simulator/physicell_reader.py
def __init__(self, default_radius: float = _DEFAULT_RADIUS):
    """
    Initialize the reader.

    Parameters
    ----------
    default_radius : float
        Radius to assign when neither ``radius`` nor ``volume`` is
        present in the snapshot. Defaults to a PhysiCell-like value
        (~8.41 micrometers, the radius of a 2494 cubic micrometer cell).
    """
    self.default_radius = float(default_radius)

load_snapshot_csv

load_snapshot_csv(csv_path: str) -> List[Dict]

Load a PhysiCell CSV snapshot.

Parameters

csv_path : str Path to a CSV file with at least the columns x, y, z, and cell_type. Optional columns: radius or volume, timestep, ID.

Returns

list of dict One dict per cell with keys x, y, z, radius, cell_type, and optionally timestep, id, volume.

Source code in tissue_simulator/physicell_reader.py
def load_snapshot_csv(self, csv_path: str) -> List[Dict]:
    """
    Load a PhysiCell CSV snapshot.

    Parameters
    ----------
    csv_path : str
        Path to a CSV file with at least the columns ``x``, ``y``, ``z``,
        and ``cell_type``. Optional columns: ``radius`` or ``volume``,
        ``timestep``, ``ID``.

    Returns
    -------
    list of dict
        One dict per cell with keys ``x``, ``y``, ``z``, ``radius``,
        ``cell_type``, and optionally ``timestep``, ``id``, ``volume``.
    """
    if not os.path.isfile(csv_path):
        raise FileNotFoundError(f"Snapshot CSV not found: {csv_path}")

    cells: List[Dict] = []
    with open(csv_path, "r", newline="") as fh:
        reader = csv.DictReader(fh)
        if reader.fieldnames is None:
            return cells

        header = list(reader.fieldnames)
        col_x = _resolve_column(header, _NUMERIC_ALIASES["x"])
        col_y = _resolve_column(header, _NUMERIC_ALIASES["y"])
        col_z = _resolve_column(header, _NUMERIC_ALIASES["z"])
        col_radius = _resolve_column(header, _NUMERIC_ALIASES["radius"])
        col_volume = _resolve_column(header, _NUMERIC_ALIASES["volume"])
        col_timestep = _resolve_column(header, _NUMERIC_ALIASES["timestep"])
        col_id = _resolve_column(header, _NUMERIC_ALIASES["id"])
        col_type = _resolve_column(header, ("cell_type", "type", "celltype"))

        if col_x is None or col_y is None or col_type is None:
            raise ValueError(
                "PhysiCell CSV must contain at minimum x, y, and "
                f"cell_type columns. Got: {header}"
            )

        for row in reader:
            x = _coerce_float(row.get(col_x))
            y = _coerce_float(row.get(col_y))
            z = _coerce_float(row.get(col_z)) if col_z else 0.0
            if x is None or y is None:
                continue
            if z is None:
                z = 0.0

            radius = None
            if col_radius is not None:
                radius = _coerce_float(row.get(col_radius))
            volume = None
            if col_volume is not None:
                volume = _coerce_float(row.get(col_volume))
            if radius is None and volume is not None:
                radius = _volume_to_radius(volume)
            if radius is None or radius <= 0:
                radius = self.default_radius

            cell_type = (row.get(col_type) or "").strip()
            if not cell_type:
                cell_type = "unknown"

            cell: Dict = {
                "x": float(x),
                "y": float(y),
                "z": float(z),
                "radius": float(radius),
                "cell_type": cell_type,
            }
            if col_volume is not None and volume is not None:
                cell["volume"] = float(volume)
            if col_timestep is not None:
                ts = _coerce_float(row.get(col_timestep))
                if ts is not None:
                    cell["timestep"] = float(ts)
            if col_id is not None:
                raw_id = row.get(col_id)
                if raw_id is not None and str(raw_id).strip() != "":
                    cell["id"] = str(raw_id).strip()

            cells.append(cell)

    return cells

load_snapshot_mat

load_snapshot_mat(mat_path: str, xml_path: Optional[str] = None) -> List[Dict]

Load a PhysiCell .mat snapshot.

Two backends are attempted, in order:

  1. pyMCDS / pcdl (optional dependency). If either of these packages is importable, the reader delegates to it for full-fidelity parsing. pyMCDS(xml_basename, output_path) is preferred because the matching output*.xml is required to recover the cell-type name mapping. When only mat_path is provided, this step is skipped (pyMCDS requires the XML).
  2. Direct scipy.io.loadmat fallback following the documented PhysiCell 1.10+ default cells-matrix layout:

=== =================== col meaning === =================== 0 ID 1 position x 2 position y 3 position z 4 total volume 5 cell type (integer) === ===================

The matrix is conventionally stored as (n_signals, n_cells), so axis 0 is the signals axis by default. We only treat axis 1 as the signals axis when axis 0 is shorter than the six documented signals and axis 1 is long enough to hold them (i.e. the file was saved transposed). If neither axis can hold the standard six signals a ValueError is raised.

Parameters

mat_path : str Path to a PhysiCell output*_cells.mat (or similar) file. xml_path : str, optional Path to the matching output*.xml. When provided, the XML is parsed for <cell_definitions> to map integer cell-type IDs to human-readable names. Without it, cell types are stringified integer IDs (e.g. "0", "1") and a warning is emitted via :mod:warnings.

Returns

list of dict Cell records in the same schema as :meth:load_snapshot_csv: keys x, y, z, radius, cell_type, and optionally volume, id.

Raises

FileNotFoundError If mat_path does not exist. ValueError If the file cannot be parsed as a MATLAB file, no cells matrix can be located, or its shape is incompatible with the documented PhysiCell layout.

Notes

The pyMCDS / pcdl package (BSD-3-Clause, https://github.com/PhysiCell-Tools/python-loader) is an optional dependency that provides reference-implementation parsing and is preferred when available. The direct fallback implements only the documented PhysiCell 1.10+ default layout and may not correctly decode custom-built PhysiCell binaries that emit a non-default signal ordering.

Source code in tissue_simulator/physicell_reader.py
def load_snapshot_mat(self,
                      mat_path: str,
                      xml_path: Optional[str] = None) -> List[Dict]:
    """
    Load a PhysiCell ``.mat`` snapshot.

    Two backends are attempted, in order:

    1. **pyMCDS / pcdl** (optional dependency). If either of these
       packages is importable, the reader delegates to it for
       full-fidelity parsing. ``pyMCDS(xml_basename, output_path)``
       is preferred because the matching ``output*.xml`` is required
       to recover the cell-type name mapping. When only ``mat_path``
       is provided, this step is skipped (pyMCDS requires the XML).
    2. **Direct ``scipy.io.loadmat`` fallback** following the
       documented PhysiCell 1.10+ default cells-matrix layout:

       ===  ===================
       col  meaning
       ===  ===================
       0    ID
       1    position x
       2    position y
       3    position z
       4    total volume
       5    cell type (integer)
       ===  ===================

       The matrix is conventionally stored as
       ``(n_signals, n_cells)``, so axis 0 is the signals axis by
       default. We only treat axis 1 as the signals axis when axis 0
       is shorter than the six documented signals and axis 1 is long
       enough to hold them (i.e. the file was saved transposed). If
       neither axis can hold the standard six signals a
       ``ValueError`` is raised.

    Parameters
    ----------
    mat_path : str
        Path to a PhysiCell ``output*_cells.mat`` (or similar) file.
    xml_path : str, optional
        Path to the matching ``output*.xml``. When provided, the XML
        is parsed for ``<cell_definitions>`` to map integer cell-type
        IDs to human-readable names. Without it, cell types are
        stringified integer IDs (e.g. ``"0"``, ``"1"``) and a warning
        is emitted via :mod:`warnings`.

    Returns
    -------
    list of dict
        Cell records in the same schema as :meth:`load_snapshot_csv`:
        keys ``x``, ``y``, ``z``, ``radius``, ``cell_type``, and
        optionally ``volume``, ``id``.

    Raises
    ------
    FileNotFoundError
        If ``mat_path`` does not exist.
    ValueError
        If the file cannot be parsed as a MATLAB file, no cells
        matrix can be located, or its shape is incompatible with the
        documented PhysiCell layout.

    Notes
    -----
    The ``pyMCDS`` / ``pcdl`` package (BSD-3-Clause,
    https://github.com/PhysiCell-Tools/python-loader) is an optional
    dependency that provides reference-implementation parsing and is
    preferred when available. The direct fallback implements only
    the documented PhysiCell 1.10+ default layout and may not
    correctly decode custom-built PhysiCell binaries that emit a
    non-default signal ordering.
    """
    if not os.path.isfile(mat_path):
        raise FileNotFoundError(f"PhysiCell .mat snapshot not found: {mat_path}")

    # 1. Try pyMCDS / pcdl first (requires xml_path because pyMCDS
    # discovers the .mat through the .xml that lives beside it).
    if xml_path is not None and os.path.isfile(xml_path):
        cells = self._try_load_mat_with_pymcds(mat_path, xml_path)
        if cells is not None:
            return cells

    # 2. Direct scipy.io.loadmat fallback.
    return self._load_mat_with_scipy(mat_path, xml_path)

load_time_series

load_time_series(output_dir: str, pattern: str = 'snapshot_*.csv') -> List[Tuple[float, List[Dict]]]

Load a directory of PhysiCell snapshot CSVs as a time series.

Parameters

output_dir : str Directory containing snapshot CSV files. pattern : str Glob pattern relative to output_dir matching snapshot files.

Returns

list of (float, list of dict) Tuples of (timestep, cells) sorted by timestep ascending. The timestep is extracted from the numeric token in each filename (e.g. snapshot_00010.csv -> 10.0).

Source code in tissue_simulator/physicell_reader.py
def load_time_series(self,
                     output_dir: str,
                     pattern: str = "snapshot_*.csv"
                     ) -> List[Tuple[float, List[Dict]]]:
    """
    Load a directory of PhysiCell snapshot CSVs as a time series.

    Parameters
    ----------
    output_dir : str
        Directory containing snapshot CSV files.
    pattern : str
        Glob pattern relative to ``output_dir`` matching snapshot files.

    Returns
    -------
    list of (float, list of dict)
        Tuples of ``(timestep, cells)`` sorted by timestep ascending.
        The timestep is extracted from the numeric token in each
        filename (e.g. ``snapshot_00010.csv`` -> 10.0).
    """
    if not os.path.isdir(output_dir):
        raise FileNotFoundError(f"Output directory not found: {output_dir}")

    matches = sorted(glob.glob(os.path.join(output_dir, pattern)))
    series: List[Tuple[float, str, List[Dict]]] = []
    for path in matches:
        timestep = _extract_timestep(path)
        cells = self.load_snapshot_csv(path)
        series.append((timestep, path, cells))

    # Sort primarily by timestep, with the filename as a stable tiebreaker
    # so collisions (e.g., files lacking digits) preserve a deterministic
    # order rather than relying on dict insertion.
    series.sort(key=lambda item: (item[0], item[1]))
    return [(ts, cells) for ts, _path, cells in series]

compute_spatial_stats

compute_spatial_stats(cells: List[Dict], mode: str = 'contact', radius_threshold: Optional[float] = None) -> Dict

Compute spatial statistics matching the tissue_simulator schema.

Parameters

cells : list of dict Cell records as returned by :meth:load_snapshot_csv. Each must contain x, y, z, radius, and cell_type. mode : str "contact" to connect cells whose surfaces touch (distance <= r_i + r_j with a 1% tolerance matching SpatialNetworkAnalyzer) or "radius" to connect cells within radius_threshold of each other (center to center). radius_threshold : float, optional Distance threshold for radius mode. Required when mode == "radius".

Returns

dict Dictionary with keys node_counts, edge_counts, and neighbor_dist.

- ``node_counts`` : ``Dict[str, int]`` mapping cell type to
  cell count.
- ``edge_counts`` : ``Dict[str, int]`` mapping a sorted
  ``"type_a-type_b"`` key to the number of contacts.
- ``neighbor_dist`` : ``Dict[str, Dict[str, float]]`` mapping
  ``type_a`` to a dict of ``type_b`` to the average number of
  ``type_b`` neighbors for a ``type_a`` cell.
Source code in tissue_simulator/physicell_reader.py
def compute_spatial_stats(self,
                          cells: List[Dict],
                          mode: str = "contact",
                          radius_threshold: Optional[float] = None) -> Dict:
    """
    Compute spatial statistics matching the ``tissue_simulator`` schema.

    Parameters
    ----------
    cells : list of dict
        Cell records as returned by :meth:`load_snapshot_csv`. Each
        must contain ``x``, ``y``, ``z``, ``radius``, and ``cell_type``.
    mode : str
        ``"contact"`` to connect cells whose surfaces touch
        (``distance <= r_i + r_j`` with a 1% tolerance matching
        ``SpatialNetworkAnalyzer``) or ``"radius"`` to connect cells
        within ``radius_threshold`` of each other (center to center).
    radius_threshold : float, optional
        Distance threshold for ``radius`` mode. Required when
        ``mode == "radius"``.

    Returns
    -------
    dict
        Dictionary with keys ``node_counts``, ``edge_counts``, and
        ``neighbor_dist``.

        - ``node_counts`` : ``Dict[str, int]`` mapping cell type to
          cell count.
        - ``edge_counts`` : ``Dict[str, int]`` mapping a sorted
          ``"type_a-type_b"`` key to the number of contacts.
        - ``neighbor_dist`` : ``Dict[str, Dict[str, float]]`` mapping
          ``type_a`` to a dict of ``type_b`` to the average number of
          ``type_b`` neighbors for a ``type_a`` cell.
    """
    if mode not in {"contact", "radius"}:
        raise ValueError(f"mode must be 'contact' or 'radius', got {mode!r}")
    if mode == "radius" and (radius_threshold is None or radius_threshold <= 0):
        raise ValueError("radius_threshold must be a positive float for 'radius' mode")
    if not SCIPY_AVAILABLE:
        raise ImportError(
            "scipy is required for spatial statistics. "
            "Install with: pip install scipy"
        )

    node_counts: Dict[str, int] = defaultdict(int)
    edge_counts: Dict[str, int] = defaultdict(int)
    neighbor_counts: Dict[str, Dict[str, int]] = defaultdict(
        lambda: defaultdict(int)
    )

    for cell in cells:
        node_counts[cell["cell_type"]] += 1

    if len(cells) < 2:
        return {
            "node_counts": dict(node_counts),
            "edge_counts": {},
            "neighbor_dist": {},
        }

    positions = np.array(
        [(c["x"], c["y"], c.get("z", 0.0)) for c in cells], dtype=float
    )
    radii = np.array([float(c["radius"]) for c in cells], dtype=float)
    types = [c["cell_type"] for c in cells]

    tree = KDTree(positions)

    if mode == "contact":
        # Tolerance matches SpatialNetworkAnalyzer (1%).
        tolerance = 1.01
        search_radius = float(radii.max()) * 2.0 * tolerance
        neighbor_lists = tree.query_ball_point(positions, r=search_radius)
        seen = set()
        for i, neighbors in enumerate(neighbor_lists):
            for j in neighbors:
                if j <= i:
                    continue
                distance = float(np.linalg.norm(positions[i] - positions[j]))
                if distance <= (radii[i] + radii[j]) * tolerance:
                    seen.add((i, j))
        edge_iter = seen
    else:
        pairs = tree.query_pairs(r=float(radius_threshold))
        edge_iter = pairs

    for i, j in edge_iter:
        t_i, t_j = types[i], types[j]
        key = "-".join(sorted([t_i, t_j]))
        edge_counts[key] += 1
        neighbor_counts[t_i][t_j] += 1
        if t_i != t_j:
            neighbor_counts[t_j][t_i] += 1
        else:
            # Both endpoints share the same type; count each direction once
            # more so the average neighbor counts reflect both endpoints.
            neighbor_counts[t_i][t_j] += 1

    neighbor_dist: Dict[str, Dict[str, float]] = {}
    for t_a, t_a_count in node_counts.items():
        if t_a_count <= 0:
            continue
        neighbor_dist[t_a] = {}
        for t_b in node_counts:
            neighbor_dist[t_a][t_b] = (
                neighbor_counts[t_a].get(t_b, 0) / t_a_count
            )

    return {
        "node_counts": dict(node_counts),
        "edge_counts": dict(edge_counts),
        "neighbor_dist": neighbor_dist,
    }

compute_stats_time_series

compute_stats_time_series(time_series: List[Tuple[float, List[Dict]]], mode: str = 'contact', radius_threshold: Optional[float] = None) -> List[Tuple[float, Dict]]

Apply :meth:compute_spatial_stats over a time series.

Parameters

time_series : list of (float, list of dict) Output of :meth:load_time_series. mode : str See :meth:compute_spatial_stats. radius_threshold : float, optional See :meth:compute_spatial_stats.

Returns

list of (float, dict) Tuples of (timestep, stats) where stats is the dict returned by :meth:compute_spatial_stats.

Source code in tissue_simulator/physicell_reader.py
def compute_stats_time_series(self,
                              time_series: List[Tuple[float, List[Dict]]],
                              mode: str = "contact",
                              radius_threshold: Optional[float] = None
                              ) -> List[Tuple[float, Dict]]:
    """
    Apply :meth:`compute_spatial_stats` over a time series.

    Parameters
    ----------
    time_series : list of (float, list of dict)
        Output of :meth:`load_time_series`.
    mode : str
        See :meth:`compute_spatial_stats`.
    radius_threshold : float, optional
        See :meth:`compute_spatial_stats`.

    Returns
    -------
    list of (float, dict)
        Tuples of ``(timestep, stats)`` where ``stats`` is the dict
        returned by :meth:`compute_spatial_stats`.
    """
    results: List[Tuple[float, Dict]] = []
    for timestep, cells in time_series:
        stats = self.compute_spatial_stats(
            cells, mode=mode, radius_threshold=radius_threshold
        )
        results.append((float(timestep), stats))
    return results

read_physicell_output

read_physicell_output(path: str, **kwargs) -> Union[List[Dict], List[Tuple[float, List[Dict]]]]

Read PhysiCell output from a file or directory.

Convenience entry point that dispatches based on path:

  • If path is a directory, returns a time series via :meth:PhysiCellReader.load_time_series.
  • If path ends in .csv, returns a single snapshot list via :meth:PhysiCellReader.load_snapshot_csv.
  • If path ends in .mat, dispatches to :meth:PhysiCellReader.load_snapshot_mat.
Parameters

path : str Path to a snapshot file or an output directory. **kwargs Forwarded to the dispatched loader (e.g. pattern for directories, xml_path for .mat files, default_radius for the reader itself).

Returns

list Either a list of cell dicts (single snapshot) or a list of (timestep, cells) tuples (directory).

Source code in tissue_simulator/physicell_reader.py
def read_physicell_output(path: str, **kwargs
                          ) -> Union[List[Dict], List[Tuple[float, List[Dict]]]]:
    """
    Read PhysiCell output from a file or directory.

    Convenience entry point that dispatches based on ``path``:

    - If ``path`` is a directory, returns a time series via
      :meth:`PhysiCellReader.load_time_series`.
    - If ``path`` ends in ``.csv``, returns a single snapshot list via
      :meth:`PhysiCellReader.load_snapshot_csv`.
    - If ``path`` ends in ``.mat``, dispatches to
      :meth:`PhysiCellReader.load_snapshot_mat`.

    Parameters
    ----------
    path : str
        Path to a snapshot file or an output directory.
    **kwargs
        Forwarded to the dispatched loader (e.g. ``pattern`` for
        directories, ``xml_path`` for ``.mat`` files,
        ``default_radius`` for the reader itself).

    Returns
    -------
    list
        Either a list of cell dicts (single snapshot) or a list of
        ``(timestep, cells)`` tuples (directory).
    """
    default_radius = kwargs.pop("default_radius", _DEFAULT_RADIUS)
    reader = PhysiCellReader(default_radius=default_radius)

    if os.path.isdir(path):
        pattern = kwargs.pop("pattern", "snapshot_*.csv")
        if kwargs:
            raise TypeError(f"Unexpected keyword arguments: {sorted(kwargs)}")
        return reader.load_time_series(path, pattern=pattern)

    lower = path.lower()
    if lower.endswith(".mat"):
        xml_path = kwargs.pop("xml_path", None)
        if kwargs:
            raise TypeError(f"Unexpected keyword arguments: {sorted(kwargs)}")
        return reader.load_snapshot_mat(path, xml_path=xml_path)

    if lower.endswith(".csv"):
        if kwargs:
            raise TypeError(f"Unexpected keyword arguments: {sorted(kwargs)}")
        return reader.load_snapshot_csv(path)

    raise ValueError(
        f"Unsupported PhysiCell output path: {path!r}. Expected a "
        "directory or a .csv/.mat file."
    )

stats_to_target_statistics

stats_to_target_statistics(stats: Dict, target_cell_count: Optional[int] = None, target_density: Optional[float] = None)

Convert a reader stats dict to a :class:TargetStatistics dataclass.

The dict returned by :meth:PhysiCellReader.compute_spatial_stats uses the schema accepted internally by :class:GraphColorizer. The :class:~tissue_simulator.replicate_generator.ReplicateGenerator expects a different shape: a list of :class:~tissue_simulator.spatial_analysis.InteractionStatistics plus cell-type proportions. This adapter performs that conversion.

Average and median distances are not present in the reader's output and are set to 0.0 on the synthesized interaction records; callers that need real distance statistics should compute them separately.

Parameters

stats : dict Output of :meth:PhysiCellReader.compute_spatial_stats. Must contain node_counts, edge_counts and neighbor_dist. target_cell_count : int, optional Optional total cell count to attach to the resulting TargetStatistics instance. When omitted, the sum of node_counts is used. target_density : float, optional Optional target packing fraction to attach.

Returns

TargetStatistics The converted dataclass, ready to pass to :class:~tissue_simulator.replicate_generator.ReplicateGenerator.

Source code in tissue_simulator/physicell_reader.py
def stats_to_target_statistics(stats: Dict,
                               target_cell_count: Optional[int] = None,
                               target_density: Optional[float] = None):
    """
    Convert a reader stats dict to a :class:`TargetStatistics` dataclass.

    The dict returned by :meth:`PhysiCellReader.compute_spatial_stats` uses
    the schema accepted internally by :class:`GraphColorizer`. The
    :class:`~tissue_simulator.replicate_generator.ReplicateGenerator` expects
    a different shape: a list of
    :class:`~tissue_simulator.spatial_analysis.InteractionStatistics` plus
    cell-type proportions. This adapter performs that conversion.

    Average and median distances are not present in the reader's output and
    are set to ``0.0`` on the synthesized interaction records; callers that
    need real distance statistics should compute them separately.

    Parameters
    ----------
    stats : dict
        Output of :meth:`PhysiCellReader.compute_spatial_stats`. Must
        contain ``node_counts``, ``edge_counts`` and ``neighbor_dist``.
    target_cell_count : int, optional
        Optional total cell count to attach to the resulting
        ``TargetStatistics`` instance. When omitted, the sum of
        ``node_counts`` is used.
    target_density : float, optional
        Optional target packing fraction to attach.

    Returns
    -------
    TargetStatistics
        The converted dataclass, ready to pass to
        :class:`~tissue_simulator.replicate_generator.ReplicateGenerator`.
    """
    # Local imports avoid a circular dependency at module load time.
    from .replicate_generator import TargetStatistics
    from .spatial_analysis import InteractionStatistics

    node_counts = dict(stats.get("node_counts") or {})
    edge_counts = dict(stats.get("edge_counts") or {})
    total_cells = sum(node_counts.values())

    cell_type_proportions: Dict[str, float] = {}
    if total_cells > 0:
        cell_type_proportions = {
            t: count / total_cells for t, count in node_counts.items()
        }

    interaction_stats: List = []
    seen_pairs = set()
    cell_types = sorted(node_counts.keys())
    for i, type_a in enumerate(cell_types):
        for type_b in cell_types[i:]:
            key = "-".join(sorted([type_a, type_b]))
            if key in seen_pairs:
                continue
            seen_pairs.add(key)
            num = int(edge_counts.get(key, 0))
            n_a = node_counts.get(type_a, 0)
            n_b = node_counts.get(type_b, 0)
            denom = (n_a * n_b) if type_a != type_b else (n_a * (n_a - 1) / 2)
            normalized = float(num) / float(denom) if denom > 0 else 0.0
            interaction_stats.append(
                InteractionStatistics(
                    type_a=type_a,
                    type_b=type_b,
                    num_interactions=num,
                    normalized_interactions=normalized,
                    avg_distance=0.0,
                    median_distance=0.0,
                )
            )

    resolved_count = (
        int(target_cell_count) if target_cell_count is not None
        else int(total_cells) if total_cells > 0 else None
    )
    return TargetStatistics(
        interaction_stats=interaction_stats,
        cell_type_proportions=cell_type_proportions or None,
        target_cell_count=resolved_count,
        target_density=target_density,
    )