Skip to content

Dynamic Simulation

This module provides the time-domain (dynamic) data generation pipeline. See the Dynamic Simulation manual page for the configuration and output reference.

Entry point

generate_dynamic_data

Generate dynamic simulation data from a YAML config. Accepted format includes: a path to the YAML file (str or os.PathLike), a dictionnary or a NestedNamespace.

Runs the full pipeline: 1. Validate config. 2. Prepare network + load scenarios. 3. Load and prepare Dynawo mappings. 4. Build solver parameters. 5. Run distributed dynamic simulations. 6. Save static (Parquet) + dynamic (Zarr) outputs.

Args

config : str | os.PathLike | dict | NestedNamespace Path to a YAML config file, a plain dict, or a NestedNamespace.

Returns

dict Paths to all generated artifacts, all rooted at settings.data_dir: the static keys (bus_data, branch_data, gen_data, y_bus_data, runtime_data, error_log, args_log, solver_log_dir, scenarios) plus dynamic_results (Zarr store) and metadata. dynamic_reports_dir is present unless dynamic.logging.save_reports is off, and final_state_values only when the variables table declares FinalStateValue rows.

Raises

TypeError If config is none of the accepted forms. ValueError If network.reader != "powsybl", the dynamic block or dynamic.dynamic_solver is missing, load.scenarios is below 1, or the removed dynamic.output_dir key is still present. RuntimeError If no sample survived, i.e. every scenario failed.

Source code in gridfm_datakit/dynamic/generate_dynamic.py
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
def generate_dynamic_data(
    config: Union[str, os.PathLike, Dict[str, Any], NestedNamespace],
) -> Dict[str, str]:
    """Generate dynamic simulation data from a YAML config.
    Accepted format includes: a path to the YAML file (str or os.PathLike), a
    dictionnary or a NestedNamespace.

    Runs the full pipeline:
    1. Validate config.
    2. Prepare network + load scenarios.
    3. Load and prepare Dynawo mappings.
    4. Build solver parameters.
    5. Run distributed dynamic simulations.
    6. Save static (Parquet) + dynamic (Zarr) outputs.

    Args
    ----
    config : str | os.PathLike | dict | NestedNamespace
        Path to a YAML config file, a plain dict, or a NestedNamespace.

    Returns
    -------
    dict
        Paths to all generated artifacts, all rooted at ``settings.data_dir``:
        the static keys (``bus_data``, ``branch_data``, ``gen_data``,
        ``y_bus_data``, ``runtime_data``, ``error_log``, ``args_log``,
        ``solver_log_dir``, ``scenarios``) plus ``dynamic_results`` (Zarr store)
        and ``metadata``. ``dynamic_reports_dir`` is present unless
        ``dynamic.logging.save_reports`` is off, and ``final_state_values`` only
        when the variables table declares FinalStateValue rows.

    Raises
    ------
    TypeError
        If ``config`` is none of the accepted forms.
    ValueError
        If ``network.reader != "powsybl"``, the ``dynamic`` block or
        ``dynamic.dynamic_solver`` is missing, ``load.scenarios`` is below 1, or
        the removed ``dynamic.output_dir`` key is still present.
    RuntimeError
        If no sample survived, i.e. every scenario failed.
    """

    # --- Step 0: load and validate config ---
    args = _load_config(config)

    _validate_dynamic_config(args)
    _configure_logging(args)

    # --- Step 1: standard environment setup (reuse generate.py logic) ---
    args, base_path, file_paths, seed = _setup_environment(args)
    # _setup_environment derives solver_log_dir (honouring enable_solver_logs)
    # into file_paths; publish it on settings so the distributed dynamic loop
    # (which reads config.settings.solver_log_dir) routes OPF + Dynawo native
    # output to files instead of dropping it.
    args.settings.solver_log_dir = file_paths["solver_log_dir"]

    # The dynamic pipeline reports progress per chunk through the
    # "gridfm_datakit.dynamic" logger, not tqdm, so _setup_environment's tqdm.log
    # stays empty and is not exported.
    #
    # It is only deleted when settings.overwrite is set: base_path has then just
    # been wiped and recreated, so the file there is certainly ours. Otherwise
    # base_path may be shared with an earlier *static* run whose accumulated
    # tqdm.log must be left alone.
    tqdm_log = file_paths.pop("tqdm_log", None)
    if tqdm_log is not None and getattr(args.settings, "overwrite", False):
        Path(tqdm_log).unlink(missing_ok=True)

    # --- Step 2: network + scenarios (reuse generate.py logic) ---
    # Only the scenarios and meta["network_path"] are used downstream: workers
    # reload the network themselves from that path (see _process_dynamic_chunk).
    _, scenarios, meta = _prepare_network_and_scenarios(args, file_paths, seed)

    # --- Step 3: dynamic inputs ---
    dynamic_inputs = load_raw_inputs(args)

    # --- Step 4: output directory ---
    # Single root: everything this run produces lives under settings.data_dir, in
    # the same base_path (data_dir/<network>/raw) the static pipeline uses for its
    # logs and scenarios. The dynamic artifacts go one level down, in dynamic/,
    # because the static pipeline writes bus_data.parquet as a *partitioned
    # directory* while we write it as a flat file: same name, different kind, so
    # they must not share a directory.
    #
    # The writer owns this directory: it recreates it from scratch, so a re-run can
    # never mix fresh artifacts with a previous run's leftovers.
    dynamic_solver = args.dynamic.dynamic_solver
    output_dir = Path(base_path) / "dynamic"

    # --- Steps 5 & 6: simulate and save, one chunk at a time ---
    # Each chunk is written and released before the next runs, so peak memory
    # tracks settings.large_chunk_size rather than the whole dataset. Dynamic
    # curves are far larger than static snapshots, which is why this streams
    # instead of collecting every sample first.
    from gridfm_datakit.dynamic.process_dynamic import iter_dynamic_simulations

    writer = _DynamicDataWriter(output_dir, file_paths, args, seed)
    for chunk_results in iter_dynamic_simulations(
        network_path=meta["network_path"],
        scenarios=scenarios,
        dynamic_inputs=dynamic_inputs,
        dynamic_solver=dynamic_solver,
        config=args,
        error_log_file=file_paths["error_log"],
        seed=seed,
    ):
        writer.write_chunk(chunk_results)
        del chunk_results
        gc.collect()
    writer.close()

    # --- Step 7: optional validation ---
    _validate_outputs(args, file_paths)

    return file_paths

Inputs

DynamicInputs

Solver-agnostic container for dynamic simulation inputs.

All three attributes are pandas DataFrames or list of pandas DataFrames so they remain compatible with pypowsybl.dynamic's native input format and are easy to inspect or serialize for debugging.

Attributes

dynamic_models : list[pd.DataFrame] List of 2 pandas DataFrames. First one is for the dynamic models that equip static elements: One row per network element to be equipped with a dynamic model. Required columns: category_name, static_id, parameter_set_id, model_name Note: unequipped static element will be given a default model Second one is for the automation systems: One row per automation system Required columns: category_name, dynamic_model_id, parameter_set_id, params, model_name events : pd.DataFrame One row per event in the simulation sequence. Required columns: event_name, static_id, start_time, params variables : pd.DataFrame One row per monitored output variables (curve or final state value). Required columns: type, model_id, variables

Source code in gridfm_datakit/dynamic/__init__.py
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
@dataclass
class DynamicInputs:
    """
    Solver-agnostic container for dynamic simulation inputs.

    All three attributes are pandas DataFrames or list of pandas DataFrames so they remain compatible with
    pypowsybl.dynamic's native input format and are easy to inspect or
    serialize for debugging.

    Attributes
    ----------
    dynamic_models : list[pd.DataFrame]
        List of 2 pandas DataFrames.
        First one is for the dynamic models that equip static elements:
            One row per network element to be equipped with a dynamic model.
            Required columns: category_name, static_id, parameter_set_id, model_name
            Note: unequipped static element will be given a default model
        Second one is for the automation systems:
            One row per automation system
            Required columns: category_name, dynamic_model_id, parameter_set_id, params, model_name
    events : pd.DataFrame
        One row per event in the simulation sequence.
        Required columns: event_name, static_id, start_time, params
    variables : pd.DataFrame
        One row per monitored output variables (curve or final state value).
        Required columns: type, model_id, variables
    """

    dynamic_models: list[pd.DataFrame]
    events: pd.DataFrame
    variables: pd.DataFrame

DynamicResults

Solver-agnostic container for dynamic simulation outputs.

Attributes

dynamic_results : pandas.DataFrame Curves for one sample, indexed by time: shape (n_timesteps, n_variables), one column per output variable. This is the solver's natural orientation and is kept as-is through the pipeline.

Note the persistent store uses the *transposed* orientation: the writer in
generate_dynamic transposes each sample to (n_variables, n_timesteps) and
stacks them into a Zarr array of shape
(n_scenarios, n_variables, n_timesteps).

report : Any Dynamic simulation report including model build-up and problem resolution. final_state_values : Any, optional Values of the variables declared as "FinalStateValue" in the variables input table: one scalar per variable at the end of the simulation, rather than a trajectory. pypowsybl returns a DataFrame indexed by the flattened variable name with a single "values" column. None when the run monitors no such variable, in which case no final_state_values.parquet is written.

Source code in gridfm_datakit/dynamic/__init__.py
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
@dataclass
class DynamicResults:
    """
    Solver-agnostic container for dynamic simulation outputs.

    Attributes
    ----------
    dynamic_results : pandas.DataFrame
        Curves for one sample, indexed by time: shape **(n_timesteps, n_variables)**,
        one column per output variable. This is the solver's natural orientation
        and is kept as-is through the pipeline.

        Note the persistent store uses the *transposed* orientation: the writer in
        generate_dynamic transposes each sample to (n_variables, n_timesteps) and
        stacks them into a Zarr array of shape
        (n_scenarios, n_variables, n_timesteps).
    report : Any
        Dynamic simulation report including model build-up and problem resolution.
    final_state_values : Any, optional
        Values of the variables declared as "FinalStateValue" in the variables
        input table: one scalar per variable at the end of the simulation, rather
        than a trajectory. pypowsybl returns a DataFrame indexed by the flattened
        variable name with a single "values" column. None when the run monitors no
        such variable, in which case no final_state_values.parquet is written.
    """

    dynamic_results: Any  # pandas.DataFrame, (n_timesteps, n_variables)
    report: Any
    final_state_values: Any = None

load_raw_inputs

Load dynamic simulation inputs from CSV files declared in the config.

Reads the four CSV (or Parquet) files listed under config.dynamic and returns a DynamicInputs instance. When dynamic_solver == "dynawo", the minimum required columns for each DataFrame are validated.

Args

args : NestedNamespace Configuration object. args.dynamic.input_files must carry: - static_element_dynamic_models_file : path to the models CSV - automation_systems_file : path to the automation systems CSV - events_file : path to the events CSV - variables_file : path to the variables CSV and args.dynamic.dynamic_solver the solver name ("dynawo" or future alternatives); it defaults to "dynawo" when absent.

Returns

DynamicInputs

Raises

FileNotFoundError If any of the four input files is missing. ValueError If required columns are absent from a DataFrame, if a key column holds an unsupported value, or if the variables table declares no "Curve" row (Dynawo solver only). TypeError If any of the input files is not of CSV or Parquet format.

Source code in gridfm_datakit/dynamic/__init__.py
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
def load_raw_inputs(
    args: NestedNamespace,
) -> DynamicInputs:
    """Load dynamic simulation inputs from CSV files declared in the config.

    Reads the four CSV (or Parquet) files listed under config.dynamic and
    returns a DynamicInputs instance. When dynamic_solver == "dynawo", the
    minimum required columns for each DataFrame are validated.

    Args
    ----
    args : NestedNamespace
        Configuration object. ``args.dynamic.input_files`` must carry:
        - static_element_dynamic_models_file     : path to the models CSV
        - automation_systems_file                : path to the automation systems CSV
        - events_file                            : path to the events CSV
        - variables_file                         : path to the variables CSV
        and ``args.dynamic.dynamic_solver`` the solver name ("dynawo" or future
        alternatives); it defaults to "dynawo" when absent.

    Returns
    -------
    DynamicInputs

    Raises
    ------
    FileNotFoundError
        If any of the four input files is missing.
    ValueError
        If required columns are absent from a DataFrame, if a key column holds an
        unsupported value, or if the variables table declares no "Curve" row
        (Dynawo solver only).
    TypeError
        If any of the input files is not of CSV or Parquet format.
    """
    dyn_input_cfg = args.dynamic.input_files

    dynamic_models = [
        _load_table(dyn_input_cfg.static_element_dynamic_models_file),
        _load_table(dyn_input_cfg.automation_systems_file),
    ]

    events = _load_table(dyn_input_cfg.events_file)
    variables = _load_table(dyn_input_cfg.variables_file)

    solver = getattr(args.dynamic, "dynamic_solver", "dynawo")
    if solver == "dynawo":
        _check_cols(
            dynamic_models[0],
            STATIC_ELEMENT_DYNAMIC_MODELS_REQUIRED_COLS,
            "static_element_dynamic_models",
        )
        _check_cols(
            dynamic_models[1],
            AUTOMATION_SYSTEMS_REQUIRED_COLS,
            "automation_systems",
        )
        _check_cols(
            events,
            EVENTS_REQUIRED_COLS,
            "events",
        )
        _check_cols(
            variables,
            VARIABLES_REQUIRED_COLS,
            "variables",
        )
        dynamic_models = [_normalize_dtypes(df) for df in dynamic_models]
        events = _normalize_dtypes(events)
        variables = _normalize_dtypes(variables)
        _validate_dynawo_values(dynamic_models[1], events, variables)

    return DynamicInputs(
        dynamic_models=dynamic_models,
        events=events,
        variables=variables,
    )

Processing

iter_dynamic_simulations

Distributed outer loop for dynamic simulation data generation.

Splits scenarios into chunks and dispatches each chunk to a worker process. Each worker initialises a Julia instance and a local copy of the pypowsybl network once, then reuses them for all scenarios in the chunk.

Yields one large chunk's samples at a time so the caller can persist and drop them: holding every sample would make peak memory scale with the whole dataset rather than with settings.large_chunk_size, and dynamic curves are far larger than static snapshots. Use process_dynamic_simulations for the collected list.

Args

network_path : str Path to the network. scenarios : np.ndarray Load scenarios array, shape (n_loads, n_scenarios, 2). dynamic_inputs: DynamicInputs Dynamics inputs. dynamic_solver : str Solver name ("dynawo" or future alternatives). config : Full NestedNamespace config. error_log_file : str Path to error log. seed : int Global seed. Deterministically derived per-chunk seeds are computed from this value.

Yields

list of dict One list per large chunk, holding one dict per successfully processed (scenario, topology-perturbation) sample, each with keys "pf_data", "dynamic_results", "scenario_index", "perturbation_index". A chunk whose scenarios all failed yields an empty list.

Source code in gridfm_datakit/dynamic/process_dynamic.py
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
def iter_dynamic_simulations(
    network_path: str,
    scenarios: np.ndarray,
    dynamic_inputs: Any,
    dynamic_solver: str,
    config: NestedNamespace,
    error_log_file: str,
    seed: int,
) -> Iterator[List[Dict[str, Any]]]:
    """Distributed outer loop for dynamic simulation data generation.

    Splits scenarios into chunks and dispatches each chunk to a worker
    process. Each worker initialises a Julia instance and a local copy of the
    pypowsybl network once, then reuses them for all scenarios in the chunk.

    Yields one large chunk's samples at a time so the caller can persist and drop
    them: holding every sample would make peak memory scale with the whole dataset
    rather than with ``settings.large_chunk_size``, and dynamic curves are far
    larger than static snapshots. Use ``process_dynamic_simulations`` for the
    collected list.

    Args
    ----
    network_path : str
        Path to the network.
    scenarios : np.ndarray
        Load scenarios array, shape (n_loads, n_scenarios, 2).
    dynamic_inputs: DynamicInputs
        Dynamics inputs.
    dynamic_solver : str
        Solver name ("dynawo" or future alternatives).
    config :
        Full NestedNamespace config.
    error_log_file : str
        Path to error log.
    seed : int
        Global seed. Deterministically derived per-chunk seeds are computed
        from this value.

    Yields
    ------
    list of dict
        One list per large chunk, holding one dict per successfully processed
        (scenario, topology-perturbation) sample, each with keys ``"pf_data"``,
        ``"dynamic_results"``, ``"scenario_index"``, ``"perturbation_index"``.
        A chunk whose scenarios all failed yields an empty list.
    """
    n_scenarios = config.load.scenarios
    large_chunk_size = config.settings.large_chunk_size
    num_processes = config.settings.num_processes
    max_iter = config.settings.max_iter
    solver_log_dir = getattr(config.settings, "solver_log_dir", None)

    # Perturbation generators, built once here (main process) from the config and
    # the base network, then passed to workers (they are picklable, mirroring the
    # static pipeline). Absent config blocks default to the identity ("none")
    # generator, so a scenario expands to exactly one sample.
    base_net = load_net(network_path).gfm_net
    _none = NestedNamespace(type="none")
    topology_generator = initialize_topology_generator(
        getattr(config, "topology_perturbation", _none),
        base_net,
    )
    # generation_perturbation randomises generation *cost* and admittance_perturbation
    # perturbs branch admittances, both pre-OPF. They therefore vary the initial
    # operating point Dynawo starts from (which machines are dispatched and at what
    # loading), and so do influence the trajectory. What they do NOT vary is the
    # dynamic model set or the event sequence: those come from the CSV inputs and are
    # identical across every sample. admittance_perturbation's r/x do reach the
    # simulated network (update_powsybl writes them), but only topology_perturbation
    # changes which elements are in service, so it alone changes the set of models
    # Dynawo instantiates and alone expands a scenario into several samples. Wired
    # for parity with the static pipeline; effect on dynamic outputs untested.
    generation_generator = initialize_generation_generator(
        getattr(config, "generation_perturbation", _none),
        base_net,
    )
    admittance_generator = initialize_admittance_generator(
        getattr(config, "admittance_perturbation", _none),
        base_net,
    )

    large_chunks = np.array_split(
        range(n_scenarios),
        int(np.ceil(n_scenarios / large_chunk_size)),
    )

    n_samples = 0

    logger.info(
        "Dynamic generation: %d scenarios in %d chunk(s), %d worker(s).",
        n_scenarios,
        len(large_chunks),
        num_processes,
    )

    for large_chunk_index, large_chunk in enumerate(large_chunks):
        chunk_size = len(large_chunk)
        scenario_chunks = np.array_split(
            large_chunk,
            min(num_processes, chunk_size),
        )

        tasks = [
            (
                chunk[0],
                chunk[-1] + 1,
                scenarios,
                network_path,
                dynamic_inputs,
                dynamic_solver,
                error_log_file,
                max_iter,
                solver_log_dir,
                seed,
                config,
                topology_generator,
                generation_generator,
                admittance_generator,
            )
            for chunk in scenario_chunks
            if len(chunk) > 0
        ]

        _mp_ctx = multiprocessing.get_context("spawn")
        with _mp_ctx.Pool(processes=num_processes) as pool:
            results = pool.map(_process_dynamic_chunk, tasks)

        samples: List[Dict[str, Any]] = []
        for chunk_results in results:
            # A worker that dies before its per-scenario loop returns [exception]
            # (see _process_dynamic_chunk). Detect both a bare exception and the
            # list-wrapped form so a failed chunk never yields non-result objects
            # (which would later crash the writer).
            if isinstance(chunk_results, Exception):
                logger.error("Error in dynamic chunk: %s", chunk_results)
            else:
                for scenario_result in chunk_results:
                    if isinstance(scenario_result, Exception):
                        logger.error("Error in dynamic chunk: %s", scenario_result)
                    else:
                        samples.append(scenario_result)

        n_samples += len(samples)
        logger.info(
            "Chunk %d/%d done (%d scenarios), %d samples so far.",
            large_chunk_index + 1,
            len(large_chunks),
            chunk_size,
            n_samples,
        )
        yield samples

process_dynamic_simulations

Collect every sample from :func:iter_dynamic_simulations into one list.

Convenience for callers that want the whole run in memory (tests, ad-hoc scripts). The pipeline streams instead; see generate_dynamic.generate_dynamic_data.

Source code in gridfm_datakit/dynamic/process_dynamic.py
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
def process_dynamic_simulations(
    network_path: str,
    scenarios: np.ndarray,
    dynamic_inputs: Any,
    dynamic_solver: str,
    config: NestedNamespace,
    error_log_file: str,
    seed: int,
) -> List[Dict[str, Any]]:
    """Collect every sample from :func:`iter_dynamic_simulations` into one list.

    Convenience for callers that want the whole run in memory (tests, ad-hoc
    scripts). The pipeline streams instead; see generate_dynamic.generate_dynamic_data.
    """
    return [
        sample
        for chunk in iter_dynamic_simulations(
            network_path=network_path,
            scenarios=scenarios,
            dynamic_inputs=dynamic_inputs,
            dynamic_solver=dynamic_solver,
            config=config,
            error_log_file=error_log_file,
            seed=seed,
        )
        for sample in chunk
    ]

process_single_dynamic_simulation

Process one load scenario, expanded over topology perturbations.

The load scenario is applied, then generation and admittance perturbations (before OPF). Each resulting topology perturbation is processed independently: the balanced initial state is computed on the perturbed network (so OPF adapts the set-points to the topology and Dynawo initialises from a converged operating point), then the dynamic simulation is run. One sample is produced per (scenario_index, perturbation_index).

Absent generators default to identity, so a scenario yields exactly one sample, the pre-perturbation behaviour.

Returns a list of result dicts (possibly empty if every perturbation failed).

Source code in gridfm_datakit/dynamic/process_dynamic.py
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
def process_single_dynamic_simulation(
    pp_net: Any,
    gfm_net: Network,
    scenarios: np.ndarray,
    scenario_index: int,
    p2g_maps,
    dynamic_mappings: Any,
    dynamic_solver_params: Any,
    dynamic_solver: str,
    julia: Any,
    topology_generator: Any = None,
    generation_generator: Any = None,
    admittance_generator: Any = None,
    error_log_file: str = None,
    lf_params: Any = None,
) -> List[Dict[str, Any]]:
    """Process one load scenario, expanded over topology perturbations.

    The load scenario is applied, then generation and admittance perturbations
    (before OPF). Each resulting topology perturbation is processed independently:
    the balanced initial state is computed on the perturbed network (so OPF adapts
    the set-points to the topology and Dynawo initialises from a converged
    operating point), then the dynamic simulation is run. One sample is produced
    per ``(scenario_index, perturbation_index)``.

    Absent generators default to identity, so a scenario yields exactly one
    sample, the pre-perturbation behaviour.

    Returns a list of result dicts (possibly empty if every perturbation failed).
    """
    # Work on a private copy so the shared per-worker network is not mutated.
    gfm_net = copy.deepcopy(gfm_net)
    gfm_net.Pd = scenarios[:, scenario_index, 0]
    gfm_net.Qd = scenarios[:, scenario_index, 1]

    # The three generators are handled differently because they have different
    # cardinality, not by accident. This mirrors the static pipeline exactly
    # (see process_network.process_scenario_*).
    #
    #   generation / admittance : 1 -> 1. They consume a stream of networks and
    #       yield one perturbed network per input, so we take next() once. They
    #       modify the operating point the OPF then optimises.
    #   topology                : 1 -> N. It yields n_topology_variants outages
    #       for a single network, so it is the only generator that expands one
    #       load scenario into several samples, hence the loop below and hence
    #       perturbation_index existing at all.
    #
    # Generation + admittance perturbations, applied before OPF.
    net_iter = iter([gfm_net])
    if generation_generator is not None:
        net_iter = generation_generator.generate(net_iter)
    if admittance_generator is not None:
        net_iter = admittance_generator.generate(net_iter)
    perturbed_base = next(net_iter)

    # Topology perturbations expand the scenario into one or more samples.
    if topology_generator is not None:
        topologies = topology_generator.generate(perturbed_base)
    else:
        topologies = [perturbed_base]

    base_variant_id = pp_net.get_working_variant_id()
    results: List[Dict[str, Any]] = []
    for perturbation_index, perturbed_net in enumerate(topologies):
        variant_id = f"scenario_{scenario_index}_perturbation_{perturbation_index}"
        variant_created = False
        try:
            # Inside the try: a failing clone_variant would otherwise escape this
            # function entirely, bypassing the handler meant to keep one bad
            # perturbation from dropping the whole scenario, and leave the working
            # variant dangling for the rest of the worker's chunk.
            pp_net.clone_variant(base_variant_id, variant_id)
            variant_created = True
            pp_net.set_working_variant(variant_id)

            # Step 1+2: balanced static state on the perturbed network
            _, pf_data = _compute_balanced_static_state(
                pp_net=pp_net,
                gfm_net=perturbed_net,
                julia=julia,
                dynamic_solver=dynamic_solver,
                p2g_maps=p2g_maps,
                scenario_index=scenario_index,
                lf_params=lf_params,
            )

            # Step 3: dynamic simulation
            dyn_results = _run_dynamic_simulation(
                pp_net,
                dynamic_mappings,
                dynamic_solver_params,
                dynamic_solver,
            )

            # Step 4: combine + label with the (scenario, perturbation) key
            combined = _combine_pf_and_dyn_res(pf_data, dyn_results)
            combined["scenario_index"] = scenario_index
            combined["perturbation_index"] = perturbation_index
            results.append(combined)
        except Exception as e:
            # A single perturbation failing must not drop the whole scenario.
            _log_error(
                error_log_file,
                f"[dynamic] scenario {scenario_index} perturbation "
                f"{perturbation_index} failed: {e}\n{traceback.format_exc()}\n",
            )
        finally:
            # Step 5: clean up the per-perturbation variant (only if it was created)
            pp_net.set_working_variant(base_variant_id)
            if variant_created:
                pp_net.remove_variant(variant_id)

    return results

Dynawo backend

DynawoMappings

Dynawo-ready simulation inputs derived from DynamicInputs. 3 mappings: - dynamic model mapping -> for both static equipments and automation systems - event mapping -> for events - variable mapping -> for the output variables

Attributes

dynamic_model_mapping : pypowsybl.dynamic.ModelMapping event_mapping : pypowsybl.dynamic.EventMapping variable_mapping : pypowsybl.dynamic.OutputVariableMapping

Source code in gridfm_datakit/dynamic/dynawo/__init__.py
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
@dataclass
class DynawoMappings:
    """
    Dynawo-ready simulation inputs derived from DynamicInputs.
    3 mappings:
        - dynamic model mapping -> for both static equipments and automation systems
        - event mapping -> for events
        - variable mapping -> for the output variables

    Attributes
    ----------
    dynamic_model_mapping : pypowsybl.dynamic.ModelMapping
    event_mapping : pypowsybl.dynamic.EventMapping
    variable_mapping : pypowsybl.dynamic.OutputVariableMapping
    """

    dynamic_model_mapping: pp.dynamic.ModelMapping
    event_mapping: pp.dynamic.EventMapping
    variable_mapping: pp.dynamic.OutputVariableMapping

generate_dynawo_mappings

Convert generic DynamicInputs into Dynawo-compatible DynawoMappings.

The conversion builds the Dynawo-specific mapping objects by parsing the generic DynamicInputs dataframes.

Args dynamic_inputs: DynamicInputs, generic inputs loaded by load_raw_inputs().

Returns DynawoMappings: simulation-ready Dynawo mappings

Source code in gridfm_datakit/dynamic/dynawo/__init__.py
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
def generate_dynawo_mappings(dynamic_inputs: DynamicInputs) -> DynawoMappings:
    """Convert generic DynamicInputs into Dynawo-compatible DynawoMappings.

    The conversion builds the Dynawo-specific mapping objects by parsing the generic
    DynamicInputs dataframes.

    Args
        dynamic_inputs: DynamicInputs, generic inputs loaded by load_raw_inputs().

    Returns
        DynawoMappings: simulation-ready Dynawo mappings
    """

    dynamic_model_mapping = _map_dynamic_models_dynawo(dynamic_inputs.dynamic_models)
    event_mapping = _map_events_dynawo(dynamic_inputs.events)
    variable_mapping = _map_variables_dynawo(dynamic_inputs.variables)

    return DynawoMappings(
        dynamic_model_mapping=dynamic_model_mapping,
        event_mapping=event_mapping,
        variable_mapping=variable_mapping,
    )

get_dynawo_simulation_parameters

Prepares the parameters for Dynawo simulation.

Raises

ValueError If a required key is missing, or an unsupported one is present.

Source code in gridfm_datakit/dynamic/dynawo/__init__.py
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
def get_dynawo_simulation_parameters(args: NestedNamespace) -> pp.dynamic.Parameters:
    """Prepares the parameters for Dynawo simulation.

    Raises
    ------
    ValueError
        If a required key is missing, or an unsupported one is present.
    """
    dict_parameters = args.dynamic.solver_parameters.to_dict()

    # This runs per worker, so a bad key fails every chunk. Both checks are what
    # make that failure name the offending setting instead of surfacing as a bare
    # KeyError from inside the provider.
    missing = sorted({"start_time", "stop_time"} - set(dict_parameters))
    if missing:
        raise ValueError(
            f"dynamic.solver_parameters: missing required key(s) {missing}. "
            "Both bound the simulation window, in seconds.",
        )

    unknown = sorted(
        set(dict_parameters)
        - set(SIMULATION_PARAMETERS_MAPPING)
        - {"start_time", "stop_time"},
    )
    if unknown:
        raise ValueError(
            f"dynamic.solver_parameters: unsupported key(s) {unknown}. Accepted: "
            f"{sorted(SIMULATION_PARAMETERS_MAPPING)} (plus start_time, stop_time).",
        )

    # provider_parameters only accept strings
    provider_parameters = {
        SIMULATION_PARAMETERS_MAPPING[key]: str(value)
        for key, value in dict_parameters.items()
        if (key not in ["start_time", "stop_time"] and value not in ["none", ""])
    }

    return pp.dynamic.Parameters(
        start_time=dict_parameters["start_time"],
        stop_time=dict_parameters["stop_time"],
        provider_parameters=provider_parameters,
    )

get_dynawo_loadflow_parameters

Build the AC load flow parameters for the balanced initial state.

The defaults deliberately differ from powsybl.get_default_lf_params(), which the static pipeline uses. Dynawo initialises each synchronous machine from the power flow solution, so the slack must sit on a machine that carries a dynamic model: slackBusSelectionMode=LARGEST_GENERATOR puts it there, and read_slack_bus/write_slack_bus keep that choice explicit and visible in the network. OpenLoadFlow's own default (MOST_MESHED) can select a bus with no generator at all, leaving Dynawo to initialise from a state its machine models cannot reproduce.

Every default is overridable through the optional dynamic.loadflow_parameters config block.

Raises

ValueError If the config block contains an unsupported key.

Source code in gridfm_datakit/dynamic/dynawo/__init__.py
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
def get_dynawo_loadflow_parameters(args: NestedNamespace) -> pp.loadflow.Parameters:
    """Build the AC load flow parameters for the balanced initial state.

    The defaults deliberately differ from ``powsybl.get_default_lf_params()``,
    which the static pipeline uses. Dynawo initialises each synchronous machine
    from the power flow solution, so the slack must sit on a machine that carries
    a dynamic model: ``slackBusSelectionMode=LARGEST_GENERATOR`` puts it there,
    and ``read_slack_bus``/``write_slack_bus`` keep that choice explicit and
    visible in the network. OpenLoadFlow's own default (MOST_MESHED) can select a
    bus with no generator at all, leaving Dynawo to initialise from a state its
    machine models cannot reproduce.

    Every default is overridable through the optional ``dynamic.loadflow_parameters``
    config block.

    Raises
    ------
    ValueError
        If the config block contains an unsupported key.
    """
    cfg = getattr(getattr(args, "dynamic", None), "loadflow_parameters", None)
    overrides = cfg.to_dict() if cfg is not None else {}

    unknown = sorted(set(overrides) - set(LOADFLOW_PARAMETERS_DEFAULTS))
    if unknown:
        raise ValueError(
            f"dynamic.loadflow_parameters: unsupported key(s) {unknown}. "
            f"Accepted: {sorted(LOADFLOW_PARAMETERS_DEFAULTS)}.",
        )

    params = {**LOADFLOW_PARAMETERS_DEFAULTS, **overrides}
    # provider_parameters is a free-form pass-through to OpenLoadFlow, which only
    # accepts strings.
    provider_parameters = {
        str(key): str(value)
        for key, value in (params["provider_parameters"] or {}).items()
    }

    return pp.loadflow.Parameters(
        distributed_slack=bool(params["distributed_slack"]),
        read_slack_bus=bool(params["read_slack_bus"]),
        write_slack_bus=bool(params["write_slack_bus"]),
        provider_parameters=provider_parameters,
    )

compute_balanced_static_state_dynawo

Compute the balanced initial conditions for a dynamic simulation.

Runs the four-step sequence required to produce a consistent initial state for Dynawo:

  1. OPF via Julia/PowerModels on the randomised gfm network → optimal dispatch for the current scenario.
  2. update_powsybl (powsybl submodule) → applies OPF results (Pg, Vm setpoints) onto the pypowsybl object with correct per-unit conventions.
  3. AC-PF via pypowsybl OpenLoadFlow → verifies convergence and produces the balanced initial state.
  4. get_pf_res / pf_post_processing (powsybl submodule) → formats pypowsybl PF results in the gridfm column schema with ID-based bus index assignment.

Args

pp_net: pypowsybl network. The caller must pass a per-worker clone/variant to avoid cross-scenario contamination gfm_net: randomised gridfm network for the current scenario (with applied load scenario and perturbations) julia: Initialised Julia interface (from "init_julia") p2g_maps: Pypowsybl-to-gridfm index maps for pp_net (from powsybl.build_p2g_maps), passed in rather than rebuilt per scenario scenario_index: int Used to label the results row (matches pf_post_processing's scenario_index argument) lf_params: pypowsybl.loadflow.Parameters for step 3 (from get_dynawo_loadflow_parameters). Defaults to the dynamic-appropriate settings in LOADFLOW_PARAMETERS_DEFAULTS when None.

Returns

pp_net: The updated pypowsybl network, balanced and ready for dynamic simulation. pf_data: dict Power flow results in gridfm column schema with keys: "bus", "gen", "branch", "Y_bus", "runtime"

Raises

RuntimeError If OPF fails to converge. ValueError If the AC power flow does not converge.

Source code in gridfm_datakit/dynamic/dynawo/simulate.py
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
def compute_balanced_static_state_dynawo(
    pp_net: pp.network.Network,
    gfm_net: Network,
    julia: Any,
    p2g_maps,
    scenario_index: int = 0,
    lf_params: Any = None,
) -> Tuple[Any, Dict[str, Any]]:
    """Compute the balanced initial conditions for a dynamic simulation.

    Runs the four-step sequence required to produce a consistent initial
    state for Dynawo:

    1. **OPF** via Julia/PowerModels on the randomised gfm network
       → optimal dispatch for the current scenario.
    2. **update_powsybl** (powsybl submodule)
       → applies OPF results (Pg, Vm setpoints) onto the pypowsybl object
       with correct per-unit conventions.
    3. **AC-PF** via pypowsybl OpenLoadFlow
       → verifies convergence and produces the balanced initial state.
    4. **get_pf_res / pf_post_processing** (powsybl submodule)
       → formats pypowsybl PF results in the gridfm column schema
       with ID-based bus index assignment.

    Args
    ----
    pp_net:
        pypowsybl network.
        The caller must pass a per-worker *clone/variant* to avoid
        cross-scenario contamination
    gfm_net:
        randomised gridfm network for the current scenario
        (with applied load scenario and perturbations)
    julia:
        Initialised Julia interface (from "init_julia")
    p2g_maps:
        Pypowsybl-to-gridfm index maps for ``pp_net`` (from
        ``powsybl.build_p2g_maps``), passed in rather than rebuilt per scenario
    scenario_index: int
        Used to label the results row (matches ``pf_post_processing``'s
        ``scenario_index`` argument)
    lf_params:
        ``pypowsybl.loadflow.Parameters`` for step 3 (from
        ``get_dynawo_loadflow_parameters``). Defaults to the dynamic-appropriate
        settings in ``LOADFLOW_PARAMETERS_DEFAULTS`` when None.

    Returns
    -------
    pp_net:
        The updated pypowsybl network, balanced and ready for dynamic
        simulation.
    pf_data: dict
        Power flow results in gridfm column schema with keys:
        ``"bus"``, ``"gen"``, ``"branch"``, ``"Y_bus"``, ``"runtime"``

    Raises
    ------
    RuntimeError
        If OPF fails to converge.
    ValueError
        If the AC power flow does not converge.
    """
    # Step 1: run OPF on the gfm network to get optimal dispatch
    opf_res = run_opf(gfm_net, julia)

    # Step 2: apply OPF setpoints to gfm_net, then push to pypowsybl
    gfm_net_pf = copy.deepcopy(gfm_net)
    gfm_net_pf = pf_preprocessing(gfm_net_pf, opf_res)

    # mapping_p2g = powsybl.build_p2g_maps(gfm_net_pf, pp_net); received as args to avoid repeated computation
    powsybl.update_powsybl(pp_net, gfm_net_pf, p2g_maps)

    # Step 3: run AC-PF via pypowsybl OpenLoadFlow.
    # These parameters deliberately differ from the static pipeline's
    # get_default_lf_params(): Dynawo initialises its synchronous machines from
    # this solution, so the slack has to sit on a machine that carries a dynamic
    # model. See get_dynawo_loadflow_parameters.
    if lf_params is None:
        from gridfm_datakit.dynamic.dynawo import get_dynawo_loadflow_parameters

        lf_params = get_dynawo_loadflow_parameters(NestedNamespace())
    t0 = time.perf_counter()
    pf_metadata = powsybl.pypowsybl.loadflow.run_ac(pp_net, lf_params)
    solve_time = time.perf_counter() - t0

    # Step 4: format results in gridfm column schema (ID-based bus assignment)
    pf_res = powsybl.get_pf_res(pp_net, solve_time, pf_metadata, p2g_maps)
    pf_data = pf_post_processing(
        scenario_index,
        gfm_net_pf,
        pf_res,
        res_dc=None,
        include_dc_res=False,
    )

    return pp_net, pf_data

run_dynawo_simulation

Apply Dynawo mappings to a balanced pypowsybl network and run the simulation.

Args

pp_net : Balanced pypowsybl network (output of compute_balanced_static_state_dynawo). dynawo_mapping : DynawoMappings Validated Dynawo-ready mappings (models, events, variables). parameters : pypowsybl.dynamic.Parameters object (from get_dynawo_simulation_parameters). drop_duplicate_timestep : Whether drop duplicate timestep in the output timeseries, True by default.

Returns

DynamicResults Solver-agnostic container holding the curves as a pandas DataFrame indexed by time with shape (n_timesteps, n_variables), plus the solver status report string and the final state values. The transpose to (n_variables, n_timesteps) happens only when writing the Zarr store.

Raises

RuntimeError If the status is not SUCCESS, or if a dynamic model was not instantiated.

Source code in gridfm_datakit/dynamic/dynawo/simulate.py
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
def run_dynawo_simulation(
    pp_net: pp.network.Network,
    dynawo_mapping: DynawoMappings,
    parameters: pp.dynamic.Parameters,
    drop_duplicate_timestep=True,
):
    """Apply Dynawo mappings to a balanced pypowsybl network and run the simulation.

    Args
    ----
    pp_net :
        Balanced pypowsybl network (output of ``compute_balanced_static_state_dynawo``).
    dynawo_mapping : DynawoMappings
        Validated Dynawo-ready mappings (models, events, variables).
    parameters :
        ``pypowsybl.dynamic.Parameters`` object (from ``get_dynawo_simulation_parameters``).
    drop_duplicate_timestep :
        Whether drop duplicate timestep in the output timeseries, True by default.

    Returns
    -------
    DynamicResults
        Solver-agnostic container holding the curves as a pandas DataFrame
        indexed by time with shape **(n_timesteps, n_variables)**, plus the solver
        status report string and the final state values. The transpose to
        (n_variables, n_timesteps) happens only when writing the Zarr store.

    Raises
    ------
    RuntimeError
        If the status is not SUCCESS, or if a dynamic model was not instantiated.
    """
    # Setup
    sim = pp.dynamic.Simulation()
    report_node = pp.report.ReportNode()

    # Run simulation. Dynawo writes prolific native output (OpenModelica banners,
    # solver iterations) straight to fd 1/2; route it through the process-wide
    # log router's "dynawo" channel so it obeys the same verbosity/file policy as
    # the OPF/PF solvers. No-op when no router is installed (e.g. ad-hoc scripts).
    with solver_capture("dynawo"):
        dyn_res = sim.run(
            pp_net,
            dynawo_mapping.dynamic_model_mapping,
            dynawo_mapping.event_mapping,
            dynawo_mapping.variable_mapping,
            parameters=parameters,
            report_node=report_node,
        )

    # Dynawo does NOT raise when the simulation fails: sim.run() returns a result
    # whose status is FAILURE and whose curves are empty or truncated. Left
    # unchecked, a diverged/aborted run would be stored as a valid trajectory (or
    # blow up the Zarr writer with a zero-width array). Raise instead: the caller
    # logs it and drops the sample, exactly like an OPF/PF divergence.
    if dyn_res.status().name != "SUCCESS":
        raise RuntimeError(
            f"Dynawo simulation failed ({dyn_res.status().name}): {dyn_res.status_text()}",
        )

    # A model Dynawo cannot instantiate is skipped, and the run still reports
    # SUCCESS: the sample would be a trajectory of a different system than the
    # input tables describe. The report is the only trace.
    report = report_node.to_json()
    failed = _failed_model_instantiations(report)
    if failed:
        raise RuntimeError(
            f"Dynawo failed to instantiate {len(failed)} dynamic model(s): {failed}. "
            "The simulation still reports SUCCESS but runs without them. Check the "
            "static_id / model_name / category_name of these rows against the "
            "network's element IDs.",
        )

    # Format results. FinalStateValue rows of the variables table are returned
    # separately from the curves and stored as a per-sample scalar table, so they
    # are carried through here rather than dropped.
    formated_dyn_res = _format_dynamic_res(dyn_res, drop_duplicate_timestep)

    return DynamicResults(
        formated_dyn_res,
        report,
        final_state_values=dyn_res.final_state_values(),
    )

Availability checks

is_dynawo_available

Return True if pypowsybl.dynamic AND a local Dynawo installation are usable.

Source code in gridfm_datakit/dynamic/dynawo/api.py
163
164
165
166
167
def is_dynawo_available() -> bool:
    """Return True if pypowsybl.dynamic AND a local Dynawo installation are usable."""
    if not is_pypowsybl_dynamic_available():
        return False
    return _dynawo_unavailable_reason() is None

check_dynawo_available

Raise if the Dynawo backend cannot run, explaining exactly what is missing.

Called before launching simulations so a missing installation surfaces as an actionable message instead of an opaque "DynawoSimulationProvider could not be instantiated" from deep inside the solver.

Raises

ImportError If pypowsybl.dynamic is not installed. RuntimeError If no usable local Dynawo installation is declared to powsybl.

Source code in gridfm_datakit/dynamic/dynawo/api.py
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
def check_dynawo_available() -> None:
    """Raise if the Dynawo backend cannot run, explaining exactly what is missing.

    Called before launching simulations so a missing installation surfaces as an
    actionable message instead of an opaque "DynawoSimulationProvider could not
    be instantiated" from deep inside the solver.

    Raises
    ------
    ImportError
        If pypowsybl.dynamic is not installed.
    RuntimeError
        If no usable local Dynawo installation is declared to powsybl.
    """
    check_pypowsybl_dynamic_available()
    reason = _dynawo_unavailable_reason()
    if reason is not None:
        raise RuntimeError(
            f"Dynawo backend unavailable: {reason}.\n\n"
            + _INSTALL_HINT.format(config_dir=_powsybl_config_dir()),
        )

Validation

validate_dynamic_data

Run the static-data validation suite on the dynamic pipeline's PF snapshot.

The dynamic pipeline stores the same physical quantities as the static one, but lays them out differently, so validate_generated_data cannot read them directly:

  • flat single-file parquet, not partitioned directories (no n_scenarios.txt);
  • a sample is keyed by the pair (scenario_index, perturbation_index), because a topology perturbation expands one load scenario into several samples, whereas the static schema has a single scenario column.

This loader bridges the two: it reads the flat files and adds a dense scenario column by ranking the distinct (scenario_index, perturbation_index) pairs, so each dynamic sample becomes one "scenario" from the checks' point of view. The checks themselves are shared verbatim with the static pipeline.

Note this validates the static snapshot (the initial operating point Dynawo starts from): the bus/branch/gen/Y-bus/runtime tables. It does not validate the time-series curves in the Zarr store.

Parameters:

Name Type Description Default
file_paths Dict[str, str]

Paths as returned by generate_dynamic_data (needs "bus_data", "branch_data", "gen_data", "y_bus_data", optionally "runtime_data").

required
mode str

Operating mode ("opf" or "pf"). The dynamic pipeline balances with an OPF but stores an AC-PF solution, so the default is "pf".

'pf'
sn_mva float

Base MVA used to scale power quantities.

100.0

Returns:

Type Description
bool

True if all validations pass.

Raises:

Type Description
AssertionError

If any validation fails.

Source code in gridfm_datakit/validation.py
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
def validate_dynamic_data(
    file_paths: Dict[str, str],
    mode: str = "pf",
    sn_mva: float = 100.0,
) -> bool:
    """Run the static-data validation suite on the dynamic pipeline's PF snapshot.

    The dynamic pipeline stores the same physical quantities as the static one, but
    lays them out differently, so validate_generated_data cannot read them directly:

    * flat single-file parquet, not partitioned directories (no ``n_scenarios.txt``);
    * a sample is keyed by the pair (scenario_index, perturbation_index), because a
      topology perturbation expands one load scenario into several samples, whereas
      the static schema has a single ``scenario`` column.

    This loader bridges the two: it reads the flat files and adds a dense
    ``scenario`` column by ranking the distinct (scenario_index, perturbation_index)
    pairs, so each dynamic sample becomes one "scenario" from the checks' point of
    view. The checks themselves are shared verbatim with the static pipeline.

    Note this validates the *static snapshot* (the initial operating point Dynawo
    starts from): the bus/branch/gen/Y-bus/runtime tables. It does not validate the
    time-series curves in the Zarr store.

    Args:
        file_paths: Paths as returned by generate_dynamic_data (needs "bus_data",
            "branch_data", "gen_data", "y_bus_data", optionally "runtime_data").
        mode: Operating mode ("opf" or "pf"). The dynamic pipeline balances with an
            OPF but stores an AC-PF solution, so the default is "pf".
        sn_mva: Base MVA used to scale power quantities.

    Returns:
        True if all validations pass.

    Raises:
        AssertionError: If any validation fails.
    """
    KEY = ["scenario_index", "perturbation_index"]

    def _read(key: str) -> pd.DataFrame:
        df = pd.read_parquet(file_paths[key], engine="pyarrow")
        # One dense "scenario" per distinct sample, consistent across every table:
        # the checks assume a single integer key and compare its set across files.
        df.insert(0, "scenario", df.set_index(KEY).index.map(sample_ids))
        return df.drop(columns=KEY)

    # Build the sample -> dense id map once, from the bus table (every table carries
    # the same set of samples), so the id is stable across all five files.
    bus_keys = pd.read_parquet(file_paths["bus_data"], columns=KEY, engine="pyarrow")
    unique_keys = bus_keys.drop_duplicates().sort_values(KEY)
    sample_ids = {
        (int(s), int(p)): i
        for i, (s, p) in enumerate(
            zip(unique_keys["scenario_index"], unique_keys["perturbation_index"]),
        )
    }

    generated_data = {
        "bus_data": _read("bus_data"),
        "branch_data": _read("branch_data"),
        "gen_data": _read("gen_data"),
        "y_bus_data": _read("y_bus_data"),
        "runtime_data": (
            _read("runtime_data") if "runtime_data" in file_paths else None
        ),
        "mode": mode,
        "file_paths": file_paths,
    }

    print(
        f"Validating dynamic static snapshot: {len(sample_ids)} samples "
        f"({len(generated_data['bus_data'])} bus rows)",
    )
    return _run_validation_checks(generated_data, mode, sn_mva)