Showing posts with label semiempirical. Show all posts
Showing posts with label semiempirical. Show all posts

Tuesday, March 6, 2018

Reviews of Random Versus Systematic Errors in Reaction Enthalpies Computed Using Semi-empirical and Minimal Basis Set Methods

We submitted this paper to ACS Omega January 31st and the reviews just came back

Reviewer: 1

Recommendation: Publish after minor revisions.

Comments:
The authors explored the CBH approach proposed by Sengupta and Raghavachari to compute the reaction enthalpy of a series of organic reactions using semi-empirical and low-cost HF/DFT as the low-level method. They also discussed the origin of errors for several cases that exhibited very large errors. The results will be very useful to the computational chemistry community, in terms of identifying effective means to compute reaction energetics and better ways to improve low-cost methods.

I have only a few minor comments:
1. It appears that dispersion correction was not included for DFTB3 and some NDDO methods. For reactions that involve very large molecules, dispersion may make a non-negligible contribution, as found, for example, for Diels-Alder reactions in the recent benchmark analysis by Gruden et al. (J. Comp. Chem. 38, 2171-2185).

2. There are several typos: line 55 of pg 2, "corrections WERE not included"; line 48 of pg 6, there is one additional "is".

3. It might be useful to report and comment on the computational cost for the different approaches. For example, PBEh-3c is still rather expensive compared to the semi-empirical methods.

4. Is there a "simple" explanation for the difference between xTB and DFTB3? For example, does the improved description of frequencies by xTB make a major difference?


Reviewer: 2

Recommendation: Publish after minor revisions.

Comments:
The authors have carried out an analysis of the performance of highly efficient computational methods for the computation of enthalpies of organic reactions using the connectivity-based hierarchy. The methods considered include DFT, HF, and a range of semi-empirical methods. The analysis is clear and some of the reported findings are indeed significant. While the good performance of DFT and HF is consistent with previous results, the lack of significant improvement with semi-empirical methods is particularly noteworthy. The paper is acceptable for publication after the authors address the following comment.

A more detailed analysis is reported for reaction 19 that is an outlier for some methods such as HF. In this system, the larger errors are attributed partly to the presence of the strained oxirane ring. Similarly, reaction 23 poses problems for some semi-empirical methods due to the presence of larger errors involving allene. In light of these observations, it may be useful to add a cautionary note to the range of problems that can be studied with such methods. I suggest a small paragraph to address the potential limitations of the inexpensive methods for such systems containing unusual bonding situations.


This work is licensed under a Creative Commons Attribution 4.0

Thursday, December 28, 2017

Predicting pKa values using PM3 - conformer search part 2

Disclaimer: These are preliminary results and may contain errors
In a previous post a looked at the effect of conformational search on the accuracy of pKa predictions.  This is a follow up from a slightly different angle using data obtained by Mads.

The plot shows 
$$ \Delta \Delta G(n) = G_{0}(n) - G_{+}(n) - (G_{0}(\min) - G_{+}(\min)) $$
where $G_{0}(n)$ is the minimum free energy value found among $n$ conformations of neutral acebutolol generated by RDKit and optimized using either PM3/COSMO (blue) or PM3/SMD (red). $ G_{+}(n)$ is the corresponding value for protonated acebutolol.  $G_{0}(n)$ (and $G_{+}(n)$) are found using $n = 11, 21, 31, ..., 201$ starting conformations and $G_{0}(\min)$ is the lowest value of  found for $G_{0}(n)$ and similarly for $G_{+}(\min)$.

So $G_{0}(\min) - G_{+}(\min)$ is our best estimate of the correct $\Delta G$ value and $ \Delta \Delta G(n) $ is the deviation from the best estimate for a give value of $n$.

Keeping in mind that a 1.4 kcal/mol error corresponds to a 1 pH unit error in pKa, we are clearly no where near convergence. The fact that we observe this using two different programs (MOPAC and GAMESS) indicates that the problem is probably not the optimizer.


The next plot shows $\Delta G_X(n) =G_X(n)-G_X(\min) $. All errors are above 1 kcal/mol for $n<50$. The error for COSMO neutral does not dip below 0.5 kcal/mol until $n =$ 181, compared to $n=$ 101 for SMD neutral.


The RDKit starting geometries are not pre-minimized using MMFF. The above plot shows the corresponding plot for MMFF optimized in the gas phase.  I think the much better convergence is due to the optimization being done in the gas phase.  It's interesting that, again, the neutral is slower to converge, and only does so for $n >$ 140. 

I think the next step is do redo this analysis with focus on finding the lowest energy conformer rather than the accuracy of the pKa values.


This work is licensed under a Creative Commons Attribution 4.0