The Danger of Relying on Queue Surveys for Model Calibration, and When is a Model Validated? (e.g. LinSig, ARCADY, PICADY)

By Simon Swanston (JCT Consultancy Ltd)
Date: 28th August 2026

Over recent years, I have become increasingly concerned that, in order to get traffic models accepted, some modellers have started adjusting model input parameters until the model queues match the queues recorded in a survey. I am also concerned that some model auditors appear to expect this to be done. The resulting model is then sometimes described as having been “validated”.

In my experience, models developed in this way can actually be some of the least fit-for-purpose models. The capacity predictions can end up being significantly different, and often much worse, than would normally be expected using conventional input parameters based on calibration and/or reasonable assumptions.

The concern is that getting the model queue to match the surveyed queue can become the main objective. Much less attention seems to be given to the more important question of why the model needs such significant adjustments in the first place.

The methodology described above is not really “validation”. Adjusting model input parameters is calibration. However, adjusting those parameters until a particular output is achieved rather misses the point of calibration. If a parameter is simply changed until the model produces the observed queue, then it is hardly surprising when the model “validates” against that queue. The validation was effectively guaranteed by the way the model was calibrated.

The problems with queue surveys

Queue surveys are not particularly reliable when it comes to accurately measuring modelled queues. In my view, they are certainly not reliable enough to be used as the main basis for calibrating a model.

There are a number of reasons for this.

  • Queues are subjective. What one person considers to be a queued vehicle, another person may consider to be a slow-moving vehicle. If several people carried out the same queue survey, there is a good chance they would record different results.
  • The survey period and model output are not necessarily measuring the same thing. Queues are often recorded in fixed intervals, such as five-minute periods, but this does not necessarily correspond with the way the modelling software calculates or reports queues.
  • Queues can change significantly over very short periods. A queue can be very different from one minute to the next, as well as from one day to another.
  • It is not always clear what is actually causing the queue. Is the queue being observed caused entirely by the junction being modelled, or are upstream or downstream conditions also contributing to it?
  • Long queues are not always measured particularly well. For example, where cameras are used, it may not be possible to identify the back of a long queue. A queue greater than 100 m may simply be recorded as “16+” vehicles.
  • Converting a queue in metres into vehicles or PCUs is also an approximation. If a PCU length of 6 m is assumed, a 120 m queue might be recorded as 20 PCUs. But drivers do not all sit 6 m apart. Some will leave considerably more space, so the actual number of vehicles in that queue could be much lower.

There is therefore quite a lot of uncertainty in a number that can sometimes be treated as though it is an accurate measurement.

The difficulty in modelling queues

Queues are inherently difficult to model accurately, regardless of the quality or level of detail used in the model.

Models simplify driver behaviour, and will not account for the significantly different decisions individual drivers may make at any given time. A junction may have the same traffic flows and signal timings on two days, yet still experience different queues, depending on the pattern of driver arrivals and decision making of drivers. Small changes in inputs, particularly when a lane approaches capacity, can result in significant differences in the resulting queue.

Both the modelled and observed queue therefore contain a degree of uncertainty. Trying to adjust a model until these two inherently variable and imperfect measurements match exactly is therefore questionable and can result in inappropriate changes to the underlying model.

ARCADY provides a good example

ARCADY modelling (i.e. JUNCTIONS11) provides a useful example of what can happen when a base model is calibrated to match a queue survey.

The standard ARCADY model derives a capacity relationship based on the arm geometry and calculates an Intercept and Slope for each arm. The Intercept represents the maximum capacity across the give-way line when opposing flow is zero, while the Slope represents the rate at which capacity reduces as opposing flow increases.

There are also various parameters that can be changed to produce a different capacity relationship. For example, a direct reduction can be applied to the derived Intercept, or an adjustment can be made to the overall capacity.

Consider an arm which is well within capacity and has a relatively low RFC. ARCADY will normally report a very low queue, potentially less than 1 PCU.

That does not necessarily mean that there is never a queue on the arm.

A platoon of vehicles could arrive together, for example. The first vehicle may stop at the give-way line, with the following vehicles queuing behind it. A few seconds later the queue may have cleared.

Someone carrying out a queue survey could therefore observe queues of 3, 4, 5 or 6 vehicles during a five-minute period, even though there may have been little or no queue for most of that five minutes.

The average queue recorded by the survey could therefore be significantly higher than the queue reported by the standard ARCADY model.

This creates a temptation to reduce the Intercept until the model queue increases sufficiently to match the survey.

The problem is that this can require a very significant reduction in capacity. The modeller may eventually find that the only way to get the model queue to match the survey is to reduce capacity to the point where the arm is close to, or even above, practical capacity. RFCs of 0.8, 0.9 or higher are not unusual in these circumstances.

Why does this happen?

At low RFCs, quite large changes in capacity can have relatively little effect on the predicted queue. For example, if capacity was halved and the RFC increased from 0.3 to 0.6, the reported queue might still only increase slightly because the arm remains well within capacity.

As the RFC gets closer to capacity, the situation changes. The predicted queue starts to increase much more rapidly. It is therefore often at this point that the modeller suddenly gets the increase in queue needed to match the survey.

But this should raise a question.

If the only way to match the surveyed queues is to make most of the arms operate at, or close to, practical capacity, is that really a sensible calibration?

Particularly if there is no obvious reason why those arms should actually be operating at such high levels of demand.

There is another point worth considering. If a queue survey recorded the queue every second rather than once every five minutes, would the calculated average queue be the same?

Probably not.

If the answer is no, then this highlights one of the problems with using an average five-minute queue survey as a calibration target. We are potentially comparing two different measurements and treating them as though they are directly comparable.

What about future-year models?

There is another, perhaps more important, issue with this approach.

If we simply adjust a model until the queues match the survey, we may never establish why the adjustments were necessary.

This matters when the model is subsequently used to assess future-year traffic flows.

If, for example, capacity has been reduced significantly to make the existing model match a surveyed queue, what gives us confidence that the same reduction is appropriate in the future year?

And when a scheme is proposed, what problem are we actually trying to solve? If the base model has effectively been forced to match a queue survey without understanding the reason for the adjustment, it becomes much harder to know whether the model is providing a fair comparison of the proposed scheme against the existing situation.

There is an even bigger problem where no queue survey exists for the proposed situation. What adjustment should be made then? And on what basis?

Is a model automatically unfit for purpose without a site visit?

The short answer is no.

There seems to be an increasing assumption that a model cannot be considered reliable unless the modeller has visited the site. I don’t think this is necessarily the case.

For a desktop study, many of the model inputs can be established from available information such as scale drawings, controller specifications and traffic surveys. An experienced modeller should also be able to make sensible assumptions about other parameters, such as saturation flows, optimised timings and unequal lane usage in ARCADY.

In fact, I would generally have much more confidence in a model based on sensible and technically defensible assumptions than one where parameters have been adjusted simply to make the queues match a survey.

After all, if a site visit is essential for a model to be reliable, what does that mean for a model of a proposed junction that does not yet exist?

That is not to say that site observations are not useful. They can be very useful, particularly for more complicated junctions.

For example, the frequency with which demand-dependent stages are called can vary from cycle to cycle. Similarly, where drivers have a choice of lanes, lane usage can vary considerably. Unless a worst-case scenario is being deliberately modelled, or it has been agreed that a demand-dependent stage is unlikely to be called, a site visit may be necessary to properly understand these things.

Site observations can also be useful where there is a disagreement between the modeller and auditor about the assumptions being made. In those circumstances, observing what actually happens at the junction can help resolve the disagreement.

The point is that a site visit should be used to improve the understanding of the junction where it is necessary, rather than simply being regarded as a prerequisite for producing a reliable model.

Summary

I think we need to be careful about falling into the trap of adjusting input parameters in LinSig, ARCADY, PICADY or other modelling software until the model queues match the average queue recorded in a survey.

Doing this can result in significant errors in modelled capacity and can leave us with a model that gives unrealistic capacity predictions.

Furthermore, if a base model was considered validated once its predicted queues matched a survey, it can result in the modeller / auditor overlooking the accuracy of individual input parameters. The focus should instead be on making sure that the model inputs are based on the best information possible. This will typically require the use of scale drawings or images, traffic surveys and controller specifications for traffic signals. Other inputs may be calibrated against site observations or set based on reasonable and technically defensible assumptions.

Queue surveys can still have an important role to play. Comparing modelled queues with surveyed queues may highlight significant differences and can be a useful way of identifying something that needs further investigation.

But a difference between the model and the survey should not automatically mean that the model input parameters need to be changed until the two numbers agree.

A queue survey should be treated as evidence to investigate, rather than a target that the model must be made to achieve.

The important question is not simply:

“Can we make the model match the observed queue?”

It is:

“Do we understand why the model and the survey are different, and is the model based on a technically reasonable representation of how the junction actually operates?”

Ultimately, it is not the purpose of this article to define when a model should be considered “validated”. That is for the modeller and, ultimately, the auditor to determine in each individual case, taking into account the requirements and procedures of the relevant Highway Authority.

What is important is that the basis for the model inputs is understood and that the modeller and auditor are satisfied that they provide a reasonable representation of how the junction operates. This is particularly important when the model is subsequently used to assess future-year scenarios, as the assumptions and input parameters need to remain appropriate.

Similarly, when modelling a proposed scheme, any changes to the model inputs should be consistent with the approach used in the base model. This helps to ensure that the base and proposed models can be fairly compared when assessing the impact of the proposal.

The intention of this article is therefore not to say what constitutes a validated model, but to highlight some of the potential pitfalls that should be considered when developing and auditing one.

No products in the basket.