Unified CRE Underwriting Engine | BGI-Integrated Specification v2.1
Document Version: 2.1 Status: Production-Ready Specification Last Updated: January 23, 2026 Classification: Institution-Grade / SR 11-7 Compliant Integration Target: BGI OS Framework (MIZ OKI 3.5+)
Executive Summary
This specification defines a compiler + factory architecture for commercial real estate (CRE) underwriting that enforces deterministic, auditable, and institution-compliant cash flow projections. The engine is designed as an 8-layer stack with:
- Canonical schemas that serve as compiler contracts
- Event contracts enabling reproducible SRPVDAL loops
- Validation gates enforced as compiler failures (not recommendations)
- Multi-tenant isolation with per-tenant run registry and KG overlay
- SR 11-7 governance embedded at every layer
Key Corrections from v2.0
| Issue | v2.0 Status | v2.1 Resolution |
|---|---|---|
| Multi-tenant assertion | Stated, not operationalized | Full Tenant Silo section with namespace boundaries |
| Canonical schema | Logic described, no contract | Appendix A: Complete object model |
| Event model | Feedback loops mentioned | Appendix B: Full event contracts |
| Layer 3 missing | Listed in stack, no body | Part IV: Complete Event Simulator spec |
| Validation gates | Described as corrections | Appendix C: Compiler failure catalog |
| Run registry | Conceptual | Appendix D: SR 11-7 audit templates |
Part I: BGI Control Plane & Multi-Tenant Enforcement
1.1 Tenant Silo Architecture
Every object, run, event, and learning update in the underwriting engine operates within strict tenant boundaries.
1.1.1 Tenant Namespace Rules
tenant_boundary_rules:
mandatory_fields:
- tenant_id: "Required on every object/run/event"
- tenant_namespace: "Logical partition key for all storage"
- tenant_environment: "prod|staging|sandbox"
isolation_enforcement:
data_isolation: "Row-level security + namespace prefixing"
compute_isolation: "Tenant-specific worker pools"
model_isolation: "No cross-tenant learning bleed"
kg_isolation: "Separate KG overlay per tenant"
1.1.2 Per-Tenant Infrastructure
| Component | Isolation Level | Implementation |
|---|---|---|
| Run Registry | Full | {tenant_id}/runs/{run_id} namespace |
| KG Overlay | Full | Separate Neo4j database per tenant |
| Model Registry | Read-shared, Write-isolated | Global templates (read-only), tenant variants (write) |
| Feature Store | Full | Tenant-partitioned BigQuery datasets |
| Audit Logs | Full | Tenant-specific Cloud Logging sinks |
| Secrets | Full | Per-tenant Secret Manager paths |
1.1.3 Cross-Tenant Protection
class TenantIsolationEnforcer:
"""
Enforces no cross-tenant data leakage at every layer.
"""
PROHIBITED_OPERATIONS = [
"cross_tenant_model_training",
"cross_tenant_feature_access",
"cross_tenant_run_comparison",
"global_learning_without_consent"
]
def validate_request(self, request: UnderwritingRequest) -> ValidationResult:
"""Every request must pass tenant boundary validation."""
if not request.tenant_id:
raise TenantBoundaryViolation("tenant_id is mandatory")
if request.references_external_tenant():
raise TenantBoundaryViolation("Cross-tenant reference prohibited")
if request.uses_global_learning() and not request.tenant_consent_flag:
raise TenantBoundaryViolation("Global learning requires explicit consent")
return ValidationResult(valid=True, tenant_id=request.tenant_id)
1.2 Data Lineage Requirements
Every high-materiality field must carry provenance metadata:
lineage_requirements:
high_materiality_fields:
- base_rent
- expense_stops
- cam_caps
- tenant_credit_rating
- cap_rate_assumptions
- market_rent_forecast
required_provenance:
doc_hash: "SHA-256 of source document"
page_number: "Page where value was extracted"
anchor_text: "Surrounding text for context"
extraction_confidence: "ML confidence score"
human_override: "Boolean + override_reason if true"
extraction_timestamp: "ISO 8601 timestamp"
1.3 Run Reproducibility Contract
Every underwriting run must be fully reproducible:
@dataclass
class RunReproducibilityContract:
"""
All inputs required to exactly reproduce an underwriting run.
"""
run_id: str # UUID
tenant_id: str # Tenant namespace
# Version pins
model_versions: Dict[str, str] # {model_name: semantic_version}
assumptions_version: str # Assumptions template version
engine_version: str # Underwriting engine version
# Data pins
data_hashes: Dict[str, str] # {data_source: SHA-256}
lease_document_hashes: List[str] # SHA-256 of each input document
# Randomness control
master_seed: int # For Monte Carlo reproducibility
# Environment
execution_timestamp: datetime
compute_environment: str # Container image hash
def verify_reproducibility(self, comparison_run: 'UnderwritingRun') -> bool:
"""Verify two runs with same contract produce identical outputs."""
return (
self.model_versions == comparison_run.contract.model_versions and
self.data_hashes == comparison_run.contract.data_hashes and
self.master_seed == comparison_run.contract.master_seed and
np.allclose(self.outputs, comparison_run.outputs, rtol=1e-9)
)
Part II: Layer 1 — Deterministic Lease Engine
2.1 Purpose
Convert lease contracts into precise, deterministic cash flow projections. The lease engine is a compiler that transforms contractual terms into executable payment schedules.
2.2 Critical Corrections (Compiler Blockers)
2.2.1 Base-Year Gross-Up Parity — BLOCKER
class GrossUpParityEnforcer:
"""
CRITICAL: Gross-up must apply IDENTICALLY to both base year
and comparison years. Asymmetric treatment is a blocker.
"""
def validate_gross_up_parity(
self,
base_year_treatment: GrossUpConfig,
comparison_year_treatment: GrossUpConfig
) -> ValidationResult:
if base_year_treatment.method != comparison_year_treatment.method:
return ValidationResult(
valid=False,
blocker=True,
error_code="GROSSUP_PARITY_001",
message=f"Base year uses {base_year_treatment.method}, "
f"comparison uses {comparison_year_treatment.method}. "
f"Methods must be identical."
)
if base_year_treatment.occupancy_threshold != comparison_year_treatment.occupancy_threshold:
return ValidationResult(
valid=False,
blocker=True,
error_code="GROSSUP_PARITY_002",
message="Occupancy thresholds must match between base and comparison years."
)
return ValidationResult(valid=True)
2.2.2 Gross-Up Sequencing — BLOCKER
gross_up_sequencing_rule:
description: "Gross-up must occur BEFORE expense stop comparison"
correct_sequence:
1: "Calculate actual building expenses"
2: "Apply gross-up to occupancy threshold"
3: "Compare grossed-up expenses to base year stop"
4: "Apply caps/limits to excess"
5: "Allocate to tenant based on pro-rata share"
blocker_violations:
- "Gross-up after stop comparison"
- "Selective gross-up (some line items only)"
- "Different gross-up timing for base vs comparison"
2.2.3 Area Classification — BLOCKER
class AreaClassificationValidator:
"""
RSF vs USF enforcement. Mixed classification is a blocker.
"""
def validate_area_consistency(self, lease: Lease) -> ValidationResult:
area_types = {
lease.base_rent_area_type,
lease.expense_allocation_area_type,
lease.cam_area_type,
lease.pro_rata_area_type
}
if len(area_types) > 1:
return ValidationResult(
valid=False,
blocker=True,
error_code="AREA_CLASS_001",
message=f"Mixed area types detected: {area_types}. "
f"All area references must use consistent RSF or USF."
)
return ValidationResult(valid=True)
2.3 Expense Classification Engine
class ExpenseClassifier:
"""
Deterministic classification of expenses as Variable, Semi-Variable, or Fixed.
"""
CLASSIFICATION_RULES = {
# Variable (scales with occupancy)
"janitorial": ("VARIABLE", 0.85), # 85% variable
"utilities_common_area": ("VARIABLE", 0.70),
"trash_removal": ("VARIABLE", 0.60),
# Semi-Variable (partial scaling)
"security": ("SEMI_VARIABLE", 0.40),
"repairs_maintenance": ("SEMI_VARIABLE", 0.50),
"management_fee": ("SEMI_VARIABLE", 0.30),
# Fixed (no occupancy scaling)
"real_estate_taxes": ("FIXED", 0.0),
"insurance": ("FIXED", 0.0),
"debt_service": ("FIXED", 0.0),
"ground_rent": ("FIXED", 0.0),
}
def classify_expense(
self,
expense_line: ExpenseLine,
lease_override: Optional[ClassificationOverride] = None
) -> ExpenseClassification:
# Lease-specific override takes precedence
if lease_override and expense_line.category in lease_override.mappings:
return lease_override.mappings[expense_line.category]
# Fall back to standard classification
if expense_line.category in self.CLASSIFICATION_RULES:
return ExpenseClassification(*self.CLASSIFICATION_RULES[expense_line.category])
# Unknown category requires human classification
raise ClassificationRequiredException(
f"Expense category '{expense_line.category}' requires manual classification"
)
2.4 Lease Cash Flow Compiler
class LeaseCashFlowCompiler:
"""
Compiles lease terms into deterministic monthly cash flows.
"""
def compile(self, lease: Lease, assumptions: AssumptionSet) -> CashFlowSchedule:
"""
Main compilation entry point. Returns fully deterministic schedule.
"""
# Validate all blockers first
self._run_blocker_validations(lease)
schedule = CashFlowSchedule(
tenant_id=lease.tenant_id,
lease_id=lease.lease_id,
start_date=lease.commencement_date,
end_date=lease.expiration_date
)
for month in lease.month_range():
monthly_cf = MonthlyCashFlow(month=month)
# Base rent (with free rent/abatements)
monthly_cf.base_rent = self._calculate_base_rent(lease, month)
# Operating expense reimbursements
monthly_cf.expense_reimbursement = self._calculate_expense_recovery(
lease, month, assumptions
)
# Additional rent items
monthly_cf.percentage_rent = self._calculate_percentage_rent(lease, month)
monthly_cf.parking_income = self._calculate_parking(lease, month)
schedule.add_month(monthly_cf)
return schedule
def _calculate_expense_recovery(
self,
lease: Lease,
month: date,
assumptions: AssumptionSet
) -> Decimal:
"""
Calculates tenant's expense recovery with correct sequencing.
"""
year = month.year
# Step 1: Get actual building expenses
actual_expenses = assumptions.get_building_expenses(year)
# Step 2: Gross-up to occupancy threshold (BEFORE stop comparison)
occupancy = assumptions.get_occupancy(year)
grossed_up = self._apply_gross_up(
actual_expenses,
occupancy,
lease.gross_up_threshold,
lease.variable_expense_classification
)
# Step 3: Get base year expenses (also grossed up with SAME rules)
base_year_expenses = self._get_grossed_up_base_year(lease)
# Step 4: Compare to stop
excess = max(Decimal("0"), grossed_up - base_year_expenses)
# Step 5: Apply caps
if lease.expense_cap:
excess = self._apply_cap(excess, lease.expense_cap, year)
# Step 6: Apply pro-rata share
tenant_share = excess * lease.pro_rata_share
return tenant_share / 12 # Monthly
Part III: Layer 2 — Tenant Credit Engine
3.1 Purpose
Model tenant default probability, recovery rates, and correlated default events using proper tail-risk methodology.
3.2 Probability of Default (PD) Framework
class TenantCreditEngine:
"""
Institution-grade tenant credit modeling with PD bands and stress multipliers.
"""
# PD bands by credit quality
PD_BANDS = {
"investment_grade": {
"AAA": (0.0001, 0.0010),
"AA": (0.0010, 0.0025),
"A": (0.0025, 0.0050),
"BBB": (0.0050, 0.0150),
},
"speculative_grade": {
"BB": (0.0150, 0.0400),
"B": (0.0400, 0.1000),
"CCC": (0.1000, 0.2500),
"CC": (0.2500, 0.4000),
"C": (0.4000, 0.6000),
}
}
# Stress multipliers for economic scenarios
STRESS_MULTIPLIERS = {
"base_case": 1.0,
"mild_recession": 1.5,
"moderate_recession": 2.5,
"severe_recession": 4.0,
"depression": 6.0,
}
def calculate_pd(
self,
tenant: Tenant,
scenario: EconomicScenario
) -> ProbabilityOfDefault:
# Get base PD from rating
base_pd = self._get_base_pd(tenant.credit_rating)
# Apply sector adjustment
sector_adj = self._get_sector_adjustment(tenant.industry_sector, scenario)
# Apply stress multiplier
stress_mult = self.STRESS_MULTIPLIERS[scenario.severity]
# Apply tenant-specific factors
tenant_adj = self._calculate_tenant_adjustment(tenant)
# Final PD (capped at 1.0)
final_pd = min(1.0, base_pd * sector_adj * stress_mult * tenant_adj)
return ProbabilityOfDefault(
tenant_id=tenant.tenant_id,
base_pd=base_pd,
adjusted_pd=final_pd,
scenario=scenario.name,
confidence_interval=self._calculate_pd_ci(final_pd)
)
3.3 Student's-t Copula for Correlated Defaults — REQUIRED
class StudentTCopulaDefaultModel:
"""
Models correlated tenant defaults using Student's-t copula.
This captures tail dependence (stress clustering) that Gaussian copulas miss.
REQUIRED: Gaussian copula is prohibited as default for CRE portfolios.
"""
def __init__(self, degrees_of_freedom: float = 4.0):
"""
Initialize with degrees of freedom for t-distribution.
Lower df = heavier tails = more stress clustering.
Typical values:
- df=4: Heavy tails, significant stress clustering
- df=8: Moderate tails
- df=30: Approaches Gaussian (not recommended for CRE)
"""
self.df = degrees_of_freedom
def simulate_correlated_defaults(
self,
tenant_pds: List[float],
correlation_matrix: np.ndarray,
n_simulations: int = 10000,
seed: int = None
) -> np.ndarray:
"""
Simulate correlated default events.
Returns: (n_simulations, n_tenants) boolean array of default events
"""
if seed is not None:
np.random.seed(seed)
n_tenants = len(tenant_pds)
# Generate correlated t-distributed random variables
# Step 1: Generate multivariate normal
mvn = np.random.multivariate_normal(
mean=np.zeros(n_tenants),
cov=correlation_matrix,
size=n_simulations
)
# Step 2: Generate chi-squared for t-distribution
chi2 = np.random.chisquare(self.df, size=(n_simulations, 1))
# Step 3: Create t-distributed samples
t_samples = mvn / np.sqrt(chi2 / self.df)
# Step 4: Transform to uniform via t-CDF
uniform = stats.t.cdf(t_samples, df=self.df)
# Step 5: Compare to PD thresholds for default determination
pd_array = np.array(tenant_pds)
defaults = uniform < pd_array
return defaults
def calculate_joint_default_probability(
self,
pd_i: float,
pd_j: float,
correlation: float
) -> float:
"""
Calculate probability of joint default between two tenants.
"""
# Transform PDs to t-quantiles
q_i = stats.t.ppf(pd_i, df=self.df)
q_j = stats.t.ppf(pd_j, df=self.df)
# Bivariate t-distribution joint probability
# (numerical integration required)
return self._bivariate_t_cdf(q_i, q_j, correlation, self.df)
3.4 Recovery Rate Model
class RecoveryRateModel:
"""
Models Loss Given Default (LGD) = 1 - Recovery Rate.
"""
# Recovery rates by lease type and market condition
RECOVERY_RATES = {
"credit_tenant": {
"normal": (0.70, 0.85), # (mean, std)
"recession": (0.55, 0.75),
"depression": (0.35, 0.55),
},
"standard": {
"normal": (0.50, 0.70),
"recession": (0.35, 0.55),
"depression": (0.20, 0.40),
},
"specialty_retail": {
"normal": (0.30, 0.50),
"recession": (0.15, 0.35),
"depression": (0.05, 0.20),
}
}
def estimate_recovery(
self,
tenant: Tenant,
market_condition: str,
remaining_term_months: int
) -> RecoveryEstimate:
base_range = self.RECOVERY_RATES[tenant.lease_type][market_condition]
# Adjust for remaining term (longer term = higher recovery)
term_adj = min(1.2, 1.0 + (remaining_term_months / 120) * 0.2)
# Adjust for market rent relationship
rent_ratio = tenant.contract_rent / tenant.market_rent
rent_adj = 1.0 if rent_ratio <= 1.0 else max(0.7, 1.0 - (rent_ratio - 1.0) * 0.3)
adjusted_mean = base_range[0] * term_adj * rent_adj
adjusted_high = base_range[1] * term_adj * rent_adj
return RecoveryEstimate(
expected_recovery=adjusted_mean,
recovery_range=(adjusted_mean * 0.8, min(1.0, adjusted_high)),
lgd=1.0 - adjusted_mean
)
Part IV: Layer 3 — Event Simulator (Discrete-Event Engine)
4.1 Purpose
Simulate discrete events that affect property cash flows: lease rollovers, tenant defaults, market shocks, and downtime periods. This layer was missing from v2.0 and is critical for proper Monte Carlo simulation.
4.2 Lease Lifecycle State Machine
class LeaseLifecycleStateMachine:
"""
Models lease states and valid transitions.
"""
STATES = {
"ACTIVE": "Lease is in force, generating cash flow",
"EXPIRING_SOON": "Within 12 months of expiration",
"EXPIRED": "Past expiration date",
"RENEWED": "Renewal executed",
"TERMINATED_EARLY": "Tenant exercised termination option",
"DEFAULTED": "Tenant in default",
"VACANT": "Space is unleased",
"IN_TI_PERIOD": "Tenant improvements underway",
"IN_FREE_RENT": "Free rent period active",
}
VALID_TRANSITIONS = {
"ACTIVE": ["EXPIRING_SOON", "TERMINATED_EARLY", "DEFAULTED"],
"EXPIRING_SOON": ["RENEWED", "EXPIRED", "DEFAULTED"],
"EXPIRED": ["VACANT", "RENEWED"],
"RENEWED": ["ACTIVE", "IN_TI_PERIOD"],
"TERMINATED_EARLY": ["VACANT"],
"DEFAULTED": ["VACANT"],
"VACANT": ["IN_TI_PERIOD", "ACTIVE"],
"IN_TI_PERIOD": ["IN_FREE_RENT", "ACTIVE"],
"IN_FREE_RENT": ["ACTIVE"],
}
def transition(
self,
current_state: str,
target_state: str,
event: LeaseEvent
) -> StateTransitionResult:
if target_state not in self.VALID_TRANSITIONS.get(current_state, []):
raise InvalidStateTransition(
f"Cannot transition from {current_state} to {target_state}"
)
return StateTransitionResult(
from_state=current_state,
to_state=target_state,
event=event,
timestamp=event.timestamp
)
4.3 Renewal Hazard Model
class RenewalHazardModel:
"""
Models probability of renewal at each point in time using hazard functions.
"""
def __init__(self):
# Base renewal probabilities by tenant type
self.base_renewal_rates = {
"credit_tenant": 0.75,
"national_chain": 0.70,
"regional_tenant": 0.60,
"local_tenant": 0.50,
"startup": 0.35,
}
def calculate_renewal_probability(
self,
tenant: Tenant,
market_conditions: MarketState,
lease: Lease
) -> float:
"""
Calculate probability tenant will renew.
"""
base_prob = self.base_renewal_rates.get(tenant.tenant_type, 0.50)
# Adjust for rent-to-market ratio
rent_ratio = lease.current_rent / market_conditions.market_rent
if rent_ratio < 0.90:
rent_adj = 1.15 # Below market = more likely to renew
elif rent_ratio > 1.10:
rent_adj = 0.75 # Above market = less likely
else:
rent_adj = 1.0
# Adjust for economic conditions
econ_adj = {
"expansion": 1.05,
"peak": 1.00,
"contraction": 0.85,
"trough": 0.75,
}.get(market_conditions.economic_regime, 1.0)
# Adjust for tenant health
health_adj = self._tenant_health_adjustment(tenant)
# Adjust for co-tenancy satisfaction
cotenancy_adj = self._cotenancy_adjustment(lease)
final_prob = base_prob * rent_adj * econ_adj * health_adj * cotenancy_adj
return min(0.95, max(0.05, final_prob))
def sample_renewal_decision(
self,
renewal_probability: float,
seed: int
) -> Tuple[bool, Optional[RenewalTerms]]:
"""
Sample whether renewal occurs and generate terms if yes.
"""
np.random.seed(seed)
renews = np.random.random() < renewal_probability
if renews:
terms = self._generate_renewal_terms()
return (True, terms)
return (False, None)
4.4 Default Event Timing Model
class DefaultTimingModel:
"""
Models WHEN a default occurs given that it will occur within the horizon.
Uses hazard rate approach for temporal distribution.
"""
def sample_default_time(
self,
annual_pd: float,
horizon_months: int,
economic_regime: str,
seed: int
) -> Optional[int]:
"""
Sample the month of default (if any) within the horizon.
Returns None if no default, otherwise the month number.
"""
np.random.seed(seed)
# Convert annual PD to monthly hazard rate
monthly_hazard = -np.log(1 - annual_pd) / 12
# Adjust hazard for economic regime timing
regime_hazard_multipliers = {
"expansion": [0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0],
"contraction": [1.0, 1.2, 1.4, 1.6, 1.8, 2.0, 2.0, 1.8, 1.6, 1.4, 1.2, 1.0],
"trough": [2.0, 2.2, 2.4, 2.2, 2.0, 1.8, 1.6, 1.4, 1.2, 1.0, 0.9, 0.8],
}
multipliers = regime_hazard_multipliers.get(
economic_regime,
[1.0] * 12
)
# Simulate month-by-month survival
for month in range(horizon_months):
monthly_mult = multipliers[month % 12]
adjusted_hazard = monthly_hazard * monthly_mult
# Check if default occurs this month
if np.random.random() < adjusted_hazard:
return month
return None # No default within horizon
4.5 Downtime and TI/LC Model
class DowntimeModel:
"""
Models vacancy duration and tenant improvement periods.
"""
# Downtime distributions by market type (months)
DOWNTIME_DISTRIBUTIONS = {
"class_a_cbd": {"mean": 4, "std": 2, "min": 1, "max": 12},
"class_a_suburban": {"mean": 6, "std": 3, "min": 2, "max": 18},
"class_b": {"mean": 9, "std": 4, "min": 3, "max": 24},
"class_c": {"mean": 14, "std": 6, "min": 6, "max": 36},
}
# TI cost ranges ($/SF)
TI_COSTS = {
"office_class_a": {"new": (60, 100), "renewal": (20, 40)},
"office_class_b": {"new": (40, 70), "renewal": (15, 30)},
"retail_inline": {"new": (30, 60), "renewal": (10, 25)},
"retail_anchor": {"new": (20, 50), "renewal": (5, 20)},
"industrial": {"new": (5, 20), "renewal": (2, 10)},
}
def sample_downtime(
self,
property_class: str,
market_conditions: MarketState,
seed: int
) -> int:
"""
Sample vacancy duration in months.
"""
np.random.seed(seed)
params = self.DOWNTIME_DISTRIBUTIONS[property_class]
# Adjust for market conditions
vacancy_adj = market_conditions.vacancy_rate / 0.10 # Normalize to 10% baseline
mean_downtime = params["mean"] * vacancy_adj
# Sample from truncated normal
downtime = np.random.normal(mean_downtime, params["std"])
downtime = int(np.clip(downtime, params["min"], params["max"]))
return downtime
def calculate_ti_lc_cost(
self,
property_type: str,
sf: float,
is_renewal: bool,
market_conditions: MarketState
) -> Decimal:
"""
Calculate tenant improvement and leasing commission costs.
"""
ti_range = self.TI_COSTS[property_type]["renewal" if is_renewal else "new"]
# Market adjustment
market_mult = 1.0 + (market_conditions.vacancy_rate - 0.10) * 2
ti_psf = (ti_range[0] + ti_range[1]) / 2 * market_mult
lc_cost = market_conditions.market_rent * 0.04 * 12 # ~4% of first year rent
total_ti = Decimal(str(ti_psf)) * Decimal(str(sf))
total_lc = Decimal(str(lc_cost)) * Decimal(str(sf))
return total_ti + total_lc
4.6 Event Timeline Generator
class EventTimelineGenerator:
"""
Generates deterministic event timelines given a seed.
Ensures simulation reproducibility.
"""
def generate_timeline(
self,
property: Property,
assumptions: AssumptionSet,
horizon_months: int,
master_seed: int
) -> EventTimeline:
"""
Generate complete event timeline for a property.
"""
timeline = EventTimeline(property_id=property.property_id)
for i, lease in enumerate(property.leases):
# Deterministic seed for each lease
lease_seed = master_seed + i * 1000
# Generate lease-specific events
events = self._generate_lease_events(
lease=lease,
assumptions=assumptions,
horizon_months=horizon_months,
seed=lease_seed
)
timeline.add_events(events)
# Sort events by timestamp
timeline.sort_by_time()
return timeline
def _generate_lease_events(
self,
lease: Lease,
assumptions: AssumptionSet,
horizon_months: int,
seed: int
) -> List[LeaseEvent]:
events = []
# Check for expiration within horizon
months_to_expiry = self._months_between(
assumptions.analysis_date,
lease.expiration_date
)
if months_to_expiry <= horizon_months:
# Generate renewal/vacancy decision
renewal_event = self._generate_renewal_event(lease, assumptions, seed)
events.append(renewal_event)
if not renewal_event.is_renewal:
# Generate downtime and re-leasing events
downtime_events = self._generate_downtime_events(
lease, assumptions, seed + 100
)
events.extend(downtime_events)
# Check for potential default
default_event = self._generate_default_event(lease, assumptions, seed + 200)
if default_event:
events.append(default_event)
return events
Part V: Layer 4 — Market Dynamics Engine
5.1 Purpose
Model market rent, vacancy, and cap rate dynamics using regime-switching mean-reverting processes. GBM is prohibited as the default model.
5.2 Regime-Switching Framework — REQUIRED
class RegimeSwitchingMarketModel:
"""
Market dynamics with discrete economic regimes.
GBM is PROHIBITED as the default - must use regime-switching mean-reversion.
"""
REGIMES = {
"expansion": {
"rent_drift": 0.03, # 3% annual growth
"vacancy_target": 0.06, # 6% target vacancy
"cap_rate_target": 0.055, # 5.5% cap rate
"volatility_mult": 0.8, # Lower volatility
},
"peak": {
"rent_drift": 0.01,
"vacancy_target": 0.05,
"cap_rate_target": 0.050,
"volatility_mult": 1.0,
},
"contraction": {
"rent_drift": -0.02,
"vacancy_target": 0.12,
"cap_rate_target": 0.070,
"volatility_mult": 1.5,
},
"trough": {
"rent_drift": -0.01,
"vacancy_target": 0.15,
"cap_rate_target": 0.085,
"volatility_mult": 1.8,
},
}
# Regime transition matrix (annual probabilities)
TRANSITION_MATRIX = np.array([
# To: Exp Peak Cont Trough
[0.70, 0.25, 0.05, 0.00], # From: Expansion
[0.10, 0.60, 0.25, 0.05], # From: Peak
[0.05, 0.10, 0.60, 0.25], # From: Contraction
[0.20, 0.05, 0.20, 0.55], # From: Trough
])
5.3 Ornstein-Uhlenbeck Exact Discretization — BLOCKER
class OUProcess:
"""
Ornstein-Uhlenbeck process with EXACT discretization.
CRITICAL: Euler discretization is a BLOCKER for mean-reverting processes.
Must use exact solution.
"""
def simulate_exact(
self,
x0: float, # Initial value
theta: float, # Mean reversion speed
mu: float, # Long-term mean
sigma: float, # Volatility
dt: float, # Time step
n_steps: int,
seed: int
) -> np.ndarray:
"""
EXACT discretization of OU process.
Euler discretization: X_{t+dt} = X_t + theta*(mu - X_t)*dt + sigma*sqrt(dt)*Z
WRONG for mean-reverting processes!
Exact discretization:
X_{t+dt} = mu + (X_t - mu)*exp(-theta*dt) + sigma*sqrt((1-exp(-2*theta*dt))/(2*theta))*Z
"""
np.random.seed(seed)
path = np.zeros(n_steps + 1)
path[0] = x0
# Pre-calculate constants for efficiency
exp_factor = np.exp(-theta * dt)
variance_factor = sigma * np.sqrt((1 - np.exp(-2 * theta * dt)) / (2 * theta))
# Generate all random innovations at once
innovations = np.random.normal(0, 1, n_steps)
for i in range(n_steps):
# EXACT discretization formula
path[i + 1] = mu + (path[i] - mu) * exp_factor + variance_factor * innovations[i]
return path
def validate_discretization_method(self, method: str) -> ValidationResult:
"""
Validation gate: ensure exact discretization is used.
"""
if method.lower() in ["euler", "euler-maruyama", "milstein"]:
return ValidationResult(
valid=False,
blocker=True,
error_code="OU_DISCRETIZATION_001",
message=f"{method} discretization is prohibited for OU processes. "
f"Must use exact discretization."
)
if method.lower() != "exact":
return ValidationResult(
valid=False,
blocker=True,
error_code="OU_DISCRETIZATION_002",
message=f"Unknown discretization method: {method}. Must be 'exact'."
)
return ValidationResult(valid=True)
5.4 Market Rent Model
class MarketRentModel:
"""
Regime-switching mean-reverting rent model.
"""
def __init__(self, market_params: MarketParameters):
self.params = market_params
self.ou_process = OUProcess()
self.regime_model = RegimeSwitchingMarketModel()
def simulate_rent_path(
self,
initial_rent: float,
initial_regime: str,
horizon_months: int,
seed: int
) -> Tuple[np.ndarray, List[str]]:
"""
Simulate market rent path with regime switches.
"""
np.random.seed(seed)
rent_path = np.zeros(horizon_months + 1)
rent_path[0] = initial_rent
regime_path = [initial_regime]
current_regime = initial_regime
for month in range(horizon_months):
# Check for regime transition (monthly probability)
if month % 12 == 0 and month > 0:
current_regime = self._sample_regime_transition(
current_regime,
seed + month
)
regime_path.append(current_regime)
# Get regime-specific parameters
regime_params = self.regime_model.REGIMES[current_regime]
# Monthly rent evolution (exact OU step)
monthly_drift = regime_params["rent_drift"] / 12
monthly_vol = self.params.rent_volatility * regime_params["volatility_mult"] / np.sqrt(12)
rent_path[month + 1] = rent_path[month] * np.exp(
monthly_drift + monthly_vol * np.random.normal()
)
return rent_path, regime_path
Part VI: Layer 5 — ML Forecast Engine
6.1 Purpose
Provide learned parameters for market dynamics, tenant behavior, and property performance. Strict leakage prevention is mandatory.
6.2 Walk-Forward Validation — BLOCKER
class WalkForwardValidator:
"""
Enforces temporal integrity for all ML models.
No information from the future can leak into training.
"""
def validate_no_leakage(
self,
training_data: pd.DataFrame,
features: List[str],
target_date_col: str,
feature_date_cols: Dict[str, str]
) -> ValidationResult:
"""
Verify no feature contains information from after the target date.
"""
violations = []
for feature in features:
feature_date_col = feature_date_cols.get(feature)
if feature_date_col:
# Check if any feature date > target date
leakage_mask = training_data[feature_date_col] > training_data[target_date_col]
if leakage_mask.any():
violations.append({
"feature": feature,
"leaking_rows": int(leakage_mask.sum()),
"example": training_data[leakage_mask].iloc[0].to_dict()
})
if violations:
return ValidationResult(
valid=False,
blocker=True,
error_code="ML_LEAKAGE_001",
message=f"Temporal leakage detected in {len(violations)} features",
details=violations
)
return ValidationResult(valid=True)
def create_walk_forward_splits(
self,
data: pd.DataFrame,
date_col: str,
min_train_periods: int = 24,
test_periods: int = 12,
step_periods: int = 6
) -> List[Tuple[pd.DataFrame, pd.DataFrame]]:
"""
Create proper walk-forward train/test splits.
"""
splits = []
dates = sorted(data[date_col].unique())
for i in range(min_train_periods, len(dates) - test_periods, step_periods):
train_end = dates[i]
test_end = dates[i + test_periods]
train = data[data[date_col] <= train_end]
test = data[(data[date_col] > train_end) & (data[date_col] <= test_end)]
splits.append((train, test))
return splits
6.3 Point-in-Time Feature Store
class PointInTimeFeatureStore:
"""
Feature store that only returns features known as of the query date.
Prevents look-ahead bias.
"""
def get_features(
self,
entity_id: str,
as_of_date: date,
feature_list: List[str]
) -> Dict[str, Any]:
"""
Retrieve features as they were known on as_of_date.
"""
features = {}
for feature_name in feature_list:
# Get the most recent value <= as_of_date
feature_history = self._get_feature_history(entity_id, feature_name)
valid_values = [
(ts, val) for ts, val in feature_history
if ts <= as_of_date
]
if valid_values:
# Use most recent valid value
_, value = max(valid_values, key=lambda x: x[0])
features[feature_name] = value
else:
features[feature_name] = None
return features
Part VII: Layer 6 — Exit/Liquidity Engine
7.1 Purpose
Model exit cap rates, refinancing risk, and time-on-market for property disposition or refinancing.
7.2 Refinancing Stress — REQUIRED When Maturity Within Horizon
class RefinancingStressModel:
"""
Models refinancing risk when debt matures within the analysis horizon.
REQUIRED validation gate when maturity is within horizon.
"""
def validate_refi_stress_required(
self,
debt: DebtStructure,
analysis_horizon: date
) -> ValidationResult:
"""
Validation gate: refi stress is required when maturity within horizon.
"""
if debt.maturity_date <= analysis_horizon:
if not self._refi_stress_applied(debt):
return ValidationResult(
valid=False,
blocker=False, # High severity, not blocker
severity="HIGH",
error_code="REFI_STRESS_001",
message=f"Debt matures on {debt.maturity_date}, within analysis horizon. "
f"Refinancing stress scenarios are required."
)
return ValidationResult(valid=True)
def generate_refi_scenarios(
self,
current_debt: DebtStructure,
property_noi: Decimal,
property_value: Decimal,
market_conditions: MarketState
) -> List[RefinancingScenario]:
"""
Generate refinancing scenarios across interest rate environments.
"""
scenarios = []
# Rate scenarios
rate_shocks = [
("base", 0.0),
("rates_up_100bps", 0.01),
("rates_up_200bps", 0.02),
("rates_up_300bps", 0.03),
("credit_crisis", 0.04),
]
for name, rate_shock in rate_shocks:
new_rate = current_debt.interest_rate + rate_shock
# Calculate max debt at new rate (DSCR constraint)
dscr_constrained_debt = self._max_debt_dscr(
property_noi, new_rate, min_dscr=1.25
)
# Calculate max debt at new rate (LTV constraint)
ltv_constrained_debt = self._max_debt_ltv(
property_value, max_ltv=0.65
)
# Binding constraint
max_new_debt = min(dscr_constrained_debt, ltv_constrained_debt)
# Equity gap
equity_gap = max(Decimal("0"), current_debt.principal - max_new_debt)
scenarios.append(RefinancingScenario(
name=name,
new_rate=new_rate,
max_debt=max_new_debt,
equity_gap=equity_gap,
dscr_at_new_debt=float(property_noi) / float(self._debt_service(max_new_debt, new_rate)),
ltv_at_new_debt=float(max_new_debt) / float(property_value),
binding_constraint="DSCR" if dscr_constrained_debt < ltv_constrained_debt else "LTV"
))
return scenarios
7.3 Time-on-Market Model
class TimeOnMarketModel:
"""
Models expected time to sell property based on market conditions.
"""
def estimate_marketing_period(
self,
property: Property,
market_conditions: MarketState,
pricing_strategy: str # "aggressive", "market", "conservative"
) -> MarketingPeriodEstimate:
base_periods = {
"class_a_core": 6,
"class_a_value_add": 9,
"class_b_core": 9,
"class_b_value_add": 12,
"class_c": 15,
"distressed": 24,
}
base_months = base_periods.get(property.investment_profile, 12)
# Market adjustment
market_mult = {
"sellers_market": 0.7,
"balanced": 1.0,
"buyers_market": 1.5,
"distressed_market": 2.5,
}.get(market_conditions.market_type, 1.0)
# Pricing adjustment
pricing_mult = {
"aggressive": 1.4,
"market": 1.0,
"conservative": 0.7,
}.get(pricing_strategy, 1.0)
expected_months = base_months * market_mult * pricing_mult
return MarketingPeriodEstimate(
expected_months=expected_months,
p10_months=expected_months * 0.5,
p90_months=expected_months * 2.0,
probability_within_6mo=self._cum_prob(6, expected_months),
probability_within_12mo=self._cum_prob(12, expected_months),
)
Part VIII: Layer 7 — Monte Carlo Engine
8.1 Purpose
Execute N-path simulations with proper convergence testing and confidence interval reporting.
8.2 Convergence Testing — REQUIRED
class MonteCarloEngine:
"""
Monte Carlo simulation with mandatory convergence testing.
"""
MINIMUM_PATHS = 10000
CONVERGENCE_THRESHOLD = 0.01 # 1% relative change threshold
def run_simulation(
self,
property: Property,
assumptions: AssumptionSet,
n_paths: int = 10000,
master_seed: int = 42
) -> SimulationResult:
# Validate minimum paths
if n_paths < self.MINIMUM_PATHS:
raise ValidationError(
f"Minimum {self.MINIMUM_PATHS} paths required, got {n_paths}"
)
results = []
for path_id in range(n_paths):
path_seed = master_seed + path_id
# Generate event timeline
timeline = self.event_generator.generate_timeline(
property, assumptions, assumptions.horizon_months, path_seed
)
# Simulate market dynamics
market_path = self.market_model.simulate(
assumptions, path_seed + 100000
)
# Calculate cash flows
cf_path = self.cash_flow_calculator.calculate(
property, timeline, market_path, assumptions
)
results.append(cf_path)
# Run convergence test
convergence = self._test_convergence(results)
return SimulationResult(
paths=results,
convergence=convergence,
summary=self._calculate_summary_statistics(results)
)
def _test_convergence(
self,
results: List[CashFlowPath]
) -> ConvergenceReport:
"""
Test for Monte Carlo convergence using running statistics.
"""
irrs = [r.irr for r in results]
# Calculate running mean at various sample sizes
sample_sizes = [1000, 2500, 5000, 7500, 10000]
running_means = []
for n in sample_sizes:
if n <= len(irrs):
running_means.append(np.mean(irrs[:n]))
# Check relative change between last two checkpoints
if len(running_means) >= 2:
relative_change = abs(
(running_means[-1] - running_means[-2]) / running_means[-2]
)
converged = relative_change < self.CONVERGENCE_THRESHOLD
else:
converged = False
relative_change = float('inf')
return ConvergenceReport(
converged=converged,
relative_change=relative_change,
sample_sizes=sample_sizes,
running_means=running_means,
recommendation="Sufficient paths" if converged else f"Consider increasing to {len(irrs) * 2} paths"
)
Part IX: Layer 8 — Feedback Loop Engine
9.1 Purpose
Capture outcomes, detect drift, and update model parameters through reinforcement learning.
9.2 Outcomes Join-Back
class OutcomesJoinbackEngine:
"""
Joins predicted outcomes with actuals for model calibration.
"""
def join_actuals(
self,
run_id: str,
prediction_month: date,
actuals: Dict[str, Any]
) -> OutcomeActual:
"""
Record actual outcomes for a previous prediction.
"""
# Retrieve original prediction
prediction = self.run_registry.get_prediction(run_id, prediction_month)
# Calculate errors
errors = {}
for metric in ['rent', 'vacancy', 'expenses', 'noi']:
if metric in prediction and metric in actuals:
errors[metric] = {
'predicted': prediction[metric],
'actual': actuals[metric],
'error': actuals[metric] - prediction[metric],
'pct_error': (actuals[metric] - prediction[metric]) / prediction[metric]
}
outcome = OutcomeActual(
run_id=run_id,
prediction_month=prediction_month,
actuals=actuals,
errors=errors,
recorded_at=datetime.utcnow()
)
# Store for drift detection
self.outcomes_store.save(outcome)
# Emit event
self.event_bus.emit(Event(
type="actuals.joined",
payload=outcome.to_dict()
))
return outcome
9.3 Drift Detection
class DriftDetector:
"""
Detects model drift by comparing prediction errors over time.
"""
def detect_drift(
self,
model_id: str,
lookback_months: int = 12,
alert_threshold: float = 0.10
) -> DriftReport:
"""
Check if model predictions are drifting from actuals.
"""
# Get recent outcomes
outcomes = self.outcomes_store.get_recent(model_id, lookback_months)
if len(outcomes) < 6:
return DriftReport(
status="insufficient_data",
message=f"Need at least 6 months of data, have {len(outcomes)}"
)
# Calculate rolling error statistics
errors = [o.errors['noi']['pct_error'] for o in outcomes if 'noi' in o.errors]
# Split into early and recent
midpoint = len(errors) // 2
early_errors = errors[:midpoint]
recent_errors = errors[midpoint:]
early_mae = np.mean(np.abs(early_errors))
recent_mae = np.mean(np.abs(recent_errors))
# Test for significant drift
drift_magnitude = recent_mae - early_mae
relative_drift = drift_magnitude / early_mae if early_mae > 0 else float('inf')
if relative_drift > alert_threshold:
status = "drift_detected"
recommendation = "Model recalibration recommended"
else:
status = "stable"
recommendation = "No action required"
report = DriftReport(
status=status,
early_mae=early_mae,
recent_mae=recent_mae,
drift_magnitude=drift_magnitude,
relative_drift=relative_drift,
recommendation=recommendation
)
if status == "drift_detected":
self.event_bus.emit(Event(
type="drift.detected",
payload=report.to_dict()
))
return report
Appendix A: Canonical Schema (Object Model)
A.1 Core Underwriting Objects
# === LEASE OBJECTS ===
@dataclass
class Lease:
"""Core lease contract representation."""
lease_id: str
tenant_id: str
property_id: str
# Term
commencement_date: date
expiration_date: date
term_months: int
# Base Rent
base_rent_schedule: List[RentStep]
rent_area_sf: Decimal
rent_area_type: Literal["RSF", "USF"]
# Expense Recovery
expense_structure: ExpenseStructure
base_year: Optional[int]
base_year_amount: Optional[Decimal]
gross_up_threshold: Decimal # e.g., 0.95 for 95%
pro_rata_share: Decimal
# Options
renewal_options: List[RenewalOption]
termination_options: List[TerminationOption]
expansion_options: List[ExpansionOption]
# Co-Tenancy
co_tenancy_clauses: List[CoTenancyClause]
# Lineage
source_document_hash: str
extraction_confidence: float
human_reviewed: bool
@dataclass
class RentStep:
"""Individual rent step in schedule."""
start_date: date
end_date: date
annual_rent: Decimal
monthly_rent: Decimal
rent_psf: Decimal
is_free_rent: bool = False
@dataclass
class ExpenseStructure:
"""Expense recovery configuration."""
structure_type: Literal["NNN", "NN", "N", "MODIFIED_GROSS", "FULL_SERVICE"]
# Expense stops
cam_stop: Optional[Decimal]
tax_stop: Optional[Decimal]
insurance_stop: Optional[Decimal]
# Caps
cam_cap_type: Optional[Literal["CUMULATIVE", "NON_CUMULATIVE", "COMPOUNDING"]]
cam_cap_percent: Optional[Decimal]
# Exclusions
excluded_expenses: List[str]
controllable_expense_cap: Optional[Decimal]
@dataclass
class Clause:
"""Generic lease clause."""
clause_id: str
clause_type: str
clause_text: str
parsed_terms: Dict[str, Any]
source_page: int
confidence: float
# === EXPENSE OBJECTS ===
@dataclass
class ExpenseLine:
"""Individual expense line item."""
category: str
sub_category: Optional[str]
amount: Decimal
year: int
classification: Literal["VARIABLE", "SEMI_VARIABLE", "FIXED"]
variable_pct: Decimal # Portion that varies with occupancy
@dataclass
class GrossUpRule:
"""Gross-up calculation rule."""
occupancy_threshold: Decimal
variable_expenses_only: bool
apply_to_base_year: bool # MUST be True
apply_to_comparison_year: bool # MUST be True
@dataclass
class BaseYear:
"""Base year expense specification."""
year: int
grossed_up_amount: Decimal
actual_amount: Decimal
occupancy_at_time: Decimal
methodology: str
# === DEBT OBJECTS ===
@dataclass
class Debt:
"""Debt structure."""
debt_id: str
principal: Decimal
interest_rate: Decimal
rate_type: Literal["FIXED", "FLOATING"]
index: Optional[str] # e.g., "SOFR"
spread: Optional[Decimal]
amortization_months: Optional[int]
maturity_date: date
io_period_months: int = 0
prepayment_type: Literal["OPEN", "DEFEASANCE", "YIELD_MAINTENANCE", "LOCKOUT"]
prepayment_end_date: Optional[date] = None
@dataclass
class RefiStress:
"""Refinancing stress scenario."""
scenario_name: str
rate_shock_bps: int
max_ltv: Decimal
min_dscr: Decimal
equity_gap: Decimal
binding_constraint: Literal["LTV", "DSCR"]
# === SIMULATION OBJECTS ===
@dataclass
class ScenarioPack:
"""Collection of scenarios for simulation."""
pack_id: str
scenarios: List[EconomicScenario]
weights: List[float] # Probability weights
correlation_matrix: np.ndarray
@dataclass
class EconomicScenario:
"""Single economic scenario."""
name: str
regime_sequence: List[str]
rent_growth_path: np.ndarray
vacancy_path: np.ndarray
cap_rate_path: np.ndarray
interest_rate_path: np.ndarray
@dataclass
class UnderwritingRun:
"""Complete underwriting run record."""
run_id: str
tenant_id: str
property_id: str
# Reproducibility contract
contract: RunReproducibilityContract
# Inputs
property: Property
assumptions: AssumptionSet
# Outputs
deterministic_cf: CashFlowSchedule
simulation_results: Optional[SimulationResult]
# Validation
validation_report: ValidationReport
# Metadata
created_at: datetime
created_by: str
status: Literal["DRAFT", "VALIDATED", "APPROVED", "SUPERSEDED"]
A.2 Validation Objects
@dataclass
class ValidationReport:
"""Complete validation report for a run."""
run_id: str
overall_status: Literal["PASS", "FAIL", "WARN"]
blocker_count: int
warning_count: int
results: List[ValidationResult]
def has_blockers(self) -> bool:
return self.blocker_count > 0
@dataclass
class ValidationResult:
"""Single validation check result."""
valid: bool
blocker: bool = False
severity: Literal["BLOCKER", "HIGH", "MEDIUM", "LOW"] = "MEDIUM"
error_code: Optional[str] = None
message: Optional[str] = None
details: Optional[Dict[str, Any]] = None
Appendix B: Event Contract (SRPVDAL Events)
B.1 Event Envelope
@dataclass
class Event:
"""Standard event envelope for all system events."""
event_id: str = field(default_factory=lambda: str(uuid.uuid4()))
event_type: str = ""
tenant_id: str = ""
timestamp: datetime = field(default_factory=datetime.utcnow)
version: str = "1.0"
source: str = "" # Which layer/service
correlation_id: Optional[str] = None # For tracing
payload: Dict[str, Any] = field(default_factory=dict)
B.2 Event Catalog
| Event Type | Source Layer | Payload | Description |
|---|---|---|---|
doc.ingested |
Ingestion | {doc_id, doc_hash, doc_type, page_count} |
Document uploaded and stored |
object.extracted |
Extraction | {object_type, object_id, confidence, source_doc} |
Object extracted from document |
object.validated |
Layer 1-6 | {object_id, validation_result, blockers} |
Object passed/failed validation |
run.created |
Orchestrator | {run_id, property_id, assumptions_version} |
New underwriting run initiated |
run.validated |
Orchestrator | {run_id, status, blocker_count} |
Run validation completed |
override.applied |
Any | {run_id, field, old_value, new_value, reason} |
Manual override applied |
scenario.executed |
Layer 7 | {run_id, scenario_id, path_count} |
Monte Carlo scenario completed |
actuals.joined |
Layer 8 | {run_id, month, errors} |
Actual outcomes joined to prediction |
drift.detected |
Layer 8 | {model_id, drift_magnitude, recommendation} |
Model drift detected |
calibration.updated |
Layer 8 | {model_id, old_params, new_params} |
Model parameters updated |
B.3 Event Contracts
class EventContracts:
"""
Schema definitions for event payloads.
"""
DOC_INGESTED = {
"required": ["doc_id", "doc_hash", "doc_type"],
"properties": {
"doc_id": {"type": "string"},
"doc_hash": {"type": "string", "pattern": "^[a-f0-9]{64}$"}, # SHA-256
"doc_type": {"type": "string", "enum": ["LEASE", "RENT_ROLL", "OPERATING_STATEMENT"]},
"page_count": {"type": "integer", "minimum": 1},
"file_size_bytes": {"type": "integer"},
}
}
OBJECT_EXTRACTED = {
"required": ["object_type", "object_id", "confidence", "source_doc"],
"properties": {
"object_type": {"type": "string"},
"object_id": {"type": "string"},
"confidence": {"type": "number", "minimum": 0, "maximum": 1},
"source_doc": {"type": "string"},
"source_page": {"type": "integer"},
"anchor_text": {"type": "string"},
}
}
RUN_CREATED = {
"required": ["run_id", "property_id", "tenant_id", "assumptions_version"],
"properties": {
"run_id": {"type": "string"},
"property_id": {"type": "string"},
"tenant_id": {"type": "string"},
"assumptions_version": {"type": "string"},
"engine_version": {"type": "string"},
"created_by": {"type": "string"},
}
}
OVERRIDE_APPLIED = {
"required": ["run_id", "field_path", "old_value", "new_value", "reason"],
"properties": {
"run_id": {"type": "string"},
"field_path": {"type": "string"}, # e.g., "lease.base_rent_schedule[0].monthly_rent"
"old_value": {}, # Any type
"new_value": {}, # Any type
"reason": {"type": "string", "minLength": 10},
"approved_by": {"type": "string"},
}
}
Appendix C: Validation Gates Catalog
C.1 Blocker Gates (Compiler Failures)
| Gate ID | Gate Name | Layer | Check | Failure Message |
|---|---|---|---|---|
GROSSUP_PARITY_001 |
Base-Year Gross-Up Parity | 1 | base_year.gross_up_method == comparison.gross_up_method |
"Gross-up methods must be identical for base and comparison years" |
GROSSUP_PARITY_002 |
Occupancy Threshold Parity | 1 | base_year.threshold == comparison.threshold |
"Occupancy thresholds must match" |
GROSSUP_SEQ_001 |
Gross-Up Sequencing | 1 | gross_up_step < stop_comparison_step |
"Gross-up must occur before stop comparison" |
AREA_CLASS_001 |
Area Classification Consistency | 1 | len(unique(area_types)) == 1 |
"All area references must use consistent RSF or USF" |
OU_DISCR_001 |
OU Exact Discretization | 4 | discretization_method == "exact" |
"Euler discretization prohibited for OU processes" |
GBM_PROHIB_001 |
GBM Prohibition | 4 | model_type != "GBM" |
"GBM prohibited as default for CRE market dynamics" |
ML_LEAK_001 |
ML Leakage Prevention | 5 | all(feature_date <= target_date) |
"Temporal leakage detected" |
WALK_FWD_001 |
Walk-Forward Validation | 5 | uses_walk_forward_splits |
"Time-series cross-validation required" |
C.2 High Severity Gates
| Gate ID | Gate Name | Layer | Check | Failure Message |
|---|---|---|---|---|
REFI_STRESS_001 |
Refinancing Stress Required | 6 | debt_maturity > horizon OR refi_stress_applied |
"Refi stress required when maturity within horizon" |
TCOPULA_001 |
Student-t Copula Default | 2 | copula_type == "student_t" |
"Gaussian copula not recommended for CRE default correlation" |
CONV_TEST_001 |
Convergence Testing | 7 | convergence_test_passed |
"Monte Carlo did not converge" |
C.3 Medium Severity Gates
| Gate ID | Gate Name | Layer | Check |
|---|---|---|---|
PD_BAND_001 |
PD Within Expected Band | 2 | PD within ±1σ of rating-implied range |
RECOV_001 |
Recovery Rate Reasonableness | 2 | Recovery rate within [0.1, 0.9] |
DOWN_001 |
Downtime Reasonableness | 3 | Downtime within market norms ±2σ |
C.4 Gate Enforcement Code
class ValidationGateRunner:
"""
Runs all validation gates and enforces blockers.
"""
def run_all_gates(
self,
run: UnderwritingRun
) -> ValidationReport:
results = []
blocker_count = 0
warning_count = 0
# Layer 1 Gates
results.extend(self._run_lease_gates(run))
# Layer 2 Gates
results.extend(self._run_credit_gates(run))
# Layer 3 Gates
results.extend(self._run_event_gates(run))
# Layer 4 Gates
results.extend(self._run_market_gates(run))
# Layer 5 Gates
results.extend(self._run_ml_gates(run))
# Layer 6 Gates
results.extend(self._run_exit_gates(run))
# Layer 7 Gates
results.extend(self._run_mc_gates(run))
for result in results:
if result.blocker:
blocker_count += 1
elif not result.valid:
warning_count += 1
overall_status = "FAIL" if blocker_count > 0 else ("WARN" if warning_count > 0 else "PASS")
return ValidationReport(
run_id=run.run_id,
overall_status=overall_status,
blocker_count=blocker_count,
warning_count=warning_count,
results=results
)
def _run_lease_gates(self, run: UnderwritingRun) -> List[ValidationResult]:
"""Run Layer 1 validation gates."""
results = []
for lease in run.property.leases:
# GROSSUP_PARITY_001
results.append(
GrossUpParityEnforcer().validate_gross_up_parity(
lease.base_year_treatment,
lease.comparison_year_treatment
)
)
# AREA_CLASS_001
results.append(
AreaClassificationValidator().validate_area_consistency(lease)
)
return results
Appendix D: Run Registry & SR 11-7 Audit Templates
D.1 Run Registry Schema
class RunRegistry:
"""
Central registry for all underwriting runs.
Provides reproducibility, audit trail, and version control.
"""
def register_run(self, run: UnderwritingRun) -> str:
"""
Register a new run with full reproducibility contract.
"""
record = {
"run_id": run.run_id,
"tenant_id": run.tenant_id,
"property_id": run.property_id,
"created_at": datetime.utcnow().isoformat(),
"created_by": run.created_by,
# Reproducibility
"contract": {
"model_versions": run.contract.model_versions,
"assumptions_version": run.contract.assumptions_version,
"engine_version": run.contract.engine_version,
"data_hashes": run.contract.data_hashes,
"master_seed": run.contract.master_seed,
"compute_environment": run.contract.compute_environment,
},
# Validation
"validation_status": run.validation_report.overall_status,
"blocker_count": run.validation_report.blocker_count,
# Outputs (references, not full data)
"output_location": f"gs://bucket/{run.tenant_id}/runs/{run.run_id}/",
# Status
"status": run.status,
}
self.store.save(record)
return run.run_id
def get_run(self, run_id: str) -> Optional[Dict]:
return self.store.get(run_id)
def list_runs(
self,
tenant_id: str,
property_id: Optional[str] = None,
status: Optional[str] = None,
limit: int = 100
) -> List[Dict]:
return self.store.query(
tenant_id=tenant_id,
property_id=property_id,
status=status,
limit=limit
)
D.2 SR 11-7 Audit Report Template
# SR 11-7 Model Validation Report
## Run Identification
- **Run ID**: {run_id}
- **Property**: {property_name} ({property_id})
- **Analysis Date**: {analysis_date}
- **Report Generated**: {report_date}
- **Generated By**: {analyst_name}
## 1. Model Development Documentation
### 1.1 Model Purpose
{model_purpose_description}
### 1.2 Model Inputs
| Input | Source | Data Quality | Validation Status |
|-------|--------|--------------|-------------------|
{input_table}
### 1.3 Model Assumptions
| Assumption | Value | Basis | Override? |
|------------|-------|-------|-----------|
{assumptions_table}
## 2. Model Validation
### 2.1 Conceptual Soundness
- [ ] Economic rationale documented
- [ ] Alternative approaches considered
- [ ] Limitations acknowledged
### 2.2 Outcomes Analysis
| Metric | Predicted | Actual | Error | Within Tolerance |
|--------|-----------|--------|-------|------------------|
{outcomes_table}
### 2.3 Validation Gate Results
| Gate ID | Gate Name | Status | Details |
|---------|-----------|--------|---------|
{gates_table}
## 3. Reproducibility Verification
### 3.1 Version Control
- Engine Version: {engine_version}
- Model Versions: {model_versions}
- Assumptions Version: {assumptions_version}
### 3.2 Data Hashes
{data_hashes}
### 3.3 Seed Value
Master Seed: {master_seed}
## 4. Overrides & Adjustments
| Field | Original Value | Override Value | Reason | Approved By |
|-------|----------------|----------------|--------|-------------|
{overrides_table}
## 5. Challenger Model Comparison
{challenger_comparison}
## 6. Conclusions & Recommendations
{conclusions}
## 7. Attestation
I attest that this model validation was performed in accordance with SR 11-7 guidance.
**Validator**: _________________ **Date**: _________
**Reviewer**: _________________ **Date**: _________
D.3 Audit Report Generator
class SR117AuditReportGenerator:
"""
Generates SR 11-7 compliant audit reports.
"""
def generate_report(
self,
run: UnderwritingRun,
outcomes: Optional[List[OutcomeActual]] = None,
challenger_runs: Optional[List[UnderwritingRun]] = None
) -> str:
"""
Generate complete SR 11-7 audit report.
"""
template = self._load_template()
# Populate sections
report = template.format(
run_id=run.run_id,
property_name=run.property.name,
property_id=run.property.property_id,
analysis_date=run.contract.execution_timestamp.date(),
report_date=date.today(),
analyst_name=run.created_by,
model_purpose_description=self._generate_purpose_section(run),
input_table=self._generate_input_table(run),
assumptions_table=self._generate_assumptions_table(run),
outcomes_table=self._generate_outcomes_table(outcomes),
gates_table=self._generate_gates_table(run.validation_report),
engine_version=run.contract.engine_version,
model_versions=json.dumps(run.contract.model_versions, indent=2),
assumptions_version=run.contract.assumptions_version,
data_hashes=json.dumps(run.contract.data_hashes, indent=2),
master_seed=run.contract.master_seed,
overrides_table=self._generate_overrides_table(run),
challenger_comparison=self._generate_challenger_section(run, challenger_runs),
conclusions=self._generate_conclusions(run),
)
return report
Appendix E: BGI Integration Checklist
E.1 Implementation Checklist
| Category | Item | Status | Notes |
|---|---|---|---|
| Multi-Tenant | Tenant ID on all objects | ☐ | |
| Per-tenant KG overlay | ☐ | ||
| Per-tenant run registry | ☐ | ||
| Cross-tenant protection enforced | ☐ | ||
| Canonical Schema | Lease object implemented | ☐ | |
| Expense objects implemented | ☐ | ||
| Debt objects implemented | ☐ | ||
| Simulation objects implemented | ☐ | ||
| Event Contracts | Event envelope standardized | ☐ | |
| All events emit to bus | ☐ | ||
| Event schemas validated | ☐ | ||
| Validation Gates | All blockers enforced | ☐ | |
| Gate runner integrated | ☐ | ||
| Blocker failures halt run | ☐ | ||
| Run Registry | Full reproducibility contract | ☐ | |
| Audit trail complete | ☐ | ||
| SR 11-7 report generator | ☐ | ||
| Layer 3 | State machine implemented | ☐ | |
| Hazard models implemented | ☐ | ||
| Event timeline generator | ☐ |
E.2 Testing Requirements
testing_requirements:
unit_tests:
- "All validation gates have unit tests"
- "All canonical objects have serialization tests"
- "Event contracts have schema validation tests"
integration_tests:
- "End-to-end run with known inputs produces expected outputs"
- "Reproducibility: same seed produces identical results"
- "Tenant isolation: cross-tenant access fails"
backtesting:
- "Walk-forward validation on historical data"
- "Out-of-sample performance within tolerance"
- "Regime transition detection accuracy"
stress_tests:
- "10,000+ path Monte Carlo completes within SLA"
- "Memory usage within bounds for large portfolios"
- "Concurrent tenant runs isolated"
Document History
| Version | Date | Author | Changes |
|---|---|---|---|
| 2.0 | Jan 2026 | — | Initial institution-grade specification |
| 2.1 | Jan 23, 2026 | BGI Integration | Added: Tenant Silo (Part I), Canonical Schema (Appendix A), Event Contracts (Appendix B), Layer 3 Event Simulator (Part IV), Validation Gates Catalog (Appendix C), Run Registry & SR 11-7 Templates (Appendix D), Fixed Part/Layer numbering |
End of Document